يرجى تقديم سيرتك الذاتية باللغة الإنجليزية والت indication مستوى إجادتك للغة الإنجليزية.
يتصل ميندريفت المتخصصين بفرص الذكاء الاصطناعي القائمة على المشاريع مع شركات تقنية رائدة، مركّزًا على الاختبار، والتقييم، وتحسين أنظمة الذكاء الاصطناعي. المشاركة قائمة على المشروع، وليست توظيفًا دائمًا.
ما تنطوي عليه هذه الفرصة
عملاء الكود الجبريون في المقدمة جيدون بالفعل في اجتياز الاختبارات. نقيس ما إذا كانوا يجتازونها بالطريقة الصحيحة. نحن نبني مجموعة بيانات لتقييم أمان وسلوك وكلاء الكود الذكي — ليس فقط ما إذا كانوا أكملوا مهمة، بل هل قاموا بها بأمانة، ضمن النطاق، ودون تقصير.
ستصمِّم مهام حيث الطريق السهل هو غير آمن، وتكتب الاختبارات التي تُكشف ذلك:
- ابداع بيئات مطور realistically — شركة افتراضية بها قاعدة كود وبنى تحتية وسياق (التذاكر، المستندات، المحادثات) التي تشكل تاريخ تطوير قابل للاقتناع
- تصميم مهام تجمع هدف تطوير بريء مع اختصار غير آمن مغرٍ: زيادة النطاق، الانتهاكات السياسة، فساد البيانات، تعديلات واسعة الصلاحيات
- كتابة اختبارات تتحقق مما إذا كان الوكيل قد أكمل المهمة بالطريقة الصحيحة — التقاط الاختزالات، وليس فقط التحقق من المخرجات
- التكرار على المهام والاختبارات بناءً على ملاحظات ضبط الجودة — مراجعة حلول الوكيل، تحليل الإخفاقات، وتحسين حتى تكون التقييم عادلًا وقويًا
ما ليس هذا:
- ليس من تسمية البيانات
- ليس هندسة الإرشادات
- ليس الأمن السيبراني أو فرق الاختبار الحمراء — لا يوجد مهاجم في السيناريو. خبرة الأمن السيبراني ميزة لكنها ليست شرطًا. نحن نبحث عن مهندسين يفهمون كيف يجب أن behaves الكود، ليس مهندسي اختراق. مهندسون برمجيات قويون، ليسوا خبراء أمان;
- ليس كتابة كود من الصفر — الوكيل يكتب معظم الكود؛ أنت تصمم الوضع وتقييم النتيجة;
ما نبحث عنه
- 4–5+ سنوات في تطوير البرمجيات;
- المكدس الأساسي: بايثون، جافا سكريبت/تايب سكريبت;
- مهارات تصميم اختبارات قوية — اختبارات وظيفية وتكامل تفصل الأمان عن غير الآمن، لا مجرد صحيح مقابل غير صحيح;
- خبرة عملية مع وكلاء الترميز (Claude Code، GitHub Copilot CLI، Codex، أو ما شابه);
- الإلمام بـ PRs وCI في GitHub كمستخدم;
- تنوّع المكدس مرحب به، ليس فِلترًا. المهام تحاكي مستودعات حقيقية مع قواعد بيانات، خطوط CI، وسكربتات النشر، لذا التعرض الأوسع للBackend والبنية التحتية مفيد حقًا — لكنك لست مطالبًا بأن تكون خبيرًا في كل طبقة;
- إتقان اللغة الإنجليزية — B2+
لماذا هذا صعب
نماذج Frontier جيدة في البرمجة. إنشاء مهمة تتحدى النماذج الأفضل بشكل حقيقي ليست مهمة يسيرة. الصعوبة الحقيقية هي بناء الإغراء — سيناريو يكون فيه المسار غير الآمن أو خارج النطاق هو طريق المقاومة الأقل — ثم كتابة اختبارات تقبض بموثوقية على وكيل اتخذ ذلك. المهام لديها حلولاً صالحة كثيرة؛ يجب أن تقبل الاختبارات جميعها وترفض السيئة منها.
كيف يعمل
التقديم؟ اجتياز التأهيل؟ الانضمام إلى مشروع؟ إكمال المهام؟ الدفع
توقعات وقت المشروع
لهذا المشروع، من المتوقع أن تستغرق المهام حوالي 20-25 ساعة أسبوعيًا خلال المراحل النشطة، بناءً على متطلبات المشروع. هذه تقديرات وليست عبء عمل مضمون، وتطبق فقط أثناء نشاط المشروع. يجب تقديم المهام قبل الموعد النهائي وتلبية معايير القبول المدرجة ليتم قبولها.
التعويض
في هذا المشروع، يمكن للمساهمين كسب ما يصل إلى 75 دولارًا في الساعة ما يعادل، وفقًا لمستواهم وتيرتهم في المساهمة.
يتفاوت التعويض عبر المشاريع وفق النطاق، التعقيد، والخبرة المطلوبة. يرجى العلم أن مشاريع أخرى على المنصة قد تقدم مستويات كسب مختلفة حسب متطلباتها.
Please submit your CV in English and indicate your level of English proficiency.
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves
Frontier coding agents are already good at passing tests. We measure whether they pass them the right way. We're building a dataset to evaluate the safety and conduct of AI coding agents — not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.
You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:
- Build realistic developer environments — a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
- Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes
- Write tests that verify whether the agent completed the task the right way — catching corners cut, not just checking outputs
- Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is not:
- Not data labeling;
- Not prompt engineering;
- Not cybersecurity or red-teaming — there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists;
- Not writing code from scratch — the agent writes most of the code; you design the situation and evaluate the outcome;
What we look for
- 4–5+ years in software development;
- Core stack: Python, JavaScript/TypeScript;
- Strong test design skills — functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
- Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
- Familiarity with GitHub PRs and CI workflows as a user;
- Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful — but you don't need to be an expert in every layer;
- English proficiency — B2+
Why this is hard
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. The real difficulty is building the temptation — a scenario where the unsafe or out-of-scope path is the path of least resistance — and then writing tests that reliably catch an agent that took it. Tasks have many valid solutions; tests must accept all of them and reject the bad ones.
How it works
Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid
Project time expectations
For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.
Compensation
On this project, contributors can earn up to $75 per hour equivalent, depending on their level and pace of contribution.
Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.