J
JobQuip
ID
العودة إلى الوظائف

AI Benchmark Engineer | Native Language Specialist - Turkish - Remote

Lilt ProductionTurkey (Remote)الراتب قابل للتفاوضعقد

الوصف الوظيفي

ABOUT THE OPPORTUNITY We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows. We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches. Note this is a remote, freelance opportunity WHAT YOU’LL DELIVER - Task Engineering: Evaluating Coding Agents. - Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling. - Prompting & Translation: finding failu

قدّم الآن

تاريخ النشر 23‏/2‏/2026
التقديم مغلق

سجّل لعرض الوظيفة كاملة والتقديم

أنشئ حسابًا مجانيًا لفتح صفحة التقديم وحفظ الوظيفة وتتبع التقدم.