Software Engineering Evaluation Specialist
Please submit your CV in English and indicate your level of English proficiency.
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
About the Role
You’ll design coding tasks that challenge frontier AI coding agents. Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome. Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable.
Responsibilities :
- Invent a realistic developer scenario — a real bug, a broken ETL, a missing feature — not a toy problem.
- Build a reproducible Docker environment with pinned dependencies.
- Write a pytest that verifies outcomes, not specific commands — deterministic, non-flaky, and does not leak the fix.
- Write an instruction.md that reads like a Jira ticket a developer would receive.
- Write a reference solve.sh proving the task is solvable.
- Calibrate difficulty so current state-of-the-art agents solve the task 20–60% of the time.
- Iterate based on feedback from expert QA reviewers.
- Later: review other authors’ tasks as a QA reviewer.
Not in scope
- Data labeling, prompt engineering.
- Production code to ship — you design problems and verification for AI agents.
- Leetcode puzzles — scenarios must look like real developer work.
- Not every candidate task ships — quality over quantity.
Requirements
- 3+ years of production software development in one backend stack — Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth.
- Python + pytest fluency — required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py.
- Docker authoring — reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.
- Linux & Bash — comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.
- AI coding agent experience — Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it.
- English — B2+ written.
Not a fit
- Data Science, ML, or Computer Vision engineers without backend-engineering output.
- Manual QA testers without automation or test authoring.
- Frontend-only, low-code / no-code, IT Support, or Business Analysts.
- Engineers who have never written pytest from scratch.
- Junior, intern, or assistant as the most recent role.
Preferred qualifications
- Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals.
- Modern Python tooling (uv, poetry, pyproject.toml).
- Coverage tooling (pytest-cov, coverage.py, gcov, llvm-cov, kcov).
- Fuzzing or property-based testing (Hypothesis).
- Prior contribution to agent-evaluation benchmarks or related frameworks.
Process
Apply → Pass qualification (90-minute sample-task screen + short behavioral interview) → Join a project → Complete tasks → Get paid.
Time commitment
- Onboarding: ~10 hours per first task.
- Steady state: ~5 hours per task, 2–4 parallel tasks per author.
- Realistic weekly load: 8–20 hours. Higher volume available for top performers.
- You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria.
Compensation:
- Paid contributions, rates up to $35/hour *.
- Task-based compensation equivalent to hourly rate, depending on performance and volume.
- Some projects include incentive payments.
*Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be provided to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project.
Apply
Submit your CV via the Mindrift platform. Indicate your English level, note this role (Software Engineering Evaluation Specialist — Terminal Bench), and include a GitHub profile link if available.
Emplois Recommandés
Repasseur (se) F/H
Tisseur de soieries pour l'ameublement, situé dans le quartier des Canuts à Lyon 4e, nous assurons une production de très haut de gamme avec les moyens techniques modernes tout en conservant une appro…
IDE/IPDE (F/H) 100% - Réanimation pédiatrique / Hôpital Femme Mère Enfant / Groupement Hospitalier Est
L'ÉTABLISSEMENT : Le Groupement Hospitalier Est regroupe l'hôpital Pierre Wertheimer, l'hôpital Louis Pradel, l'hôpital Femme Mère Enfant et l'Institut d'Hématologie et d'Oncologie Pédiatrique et as…
Paysagiste Chargé.e de projet / Agence MOZ Lyon
MOZ Paysage réunit architectes, paysagistes et urbanistes, mettant leur complémentarité au service d’une approche nouvelle du paysage dans une démarche de co-construction. Rayonnant sur le quart Sud…
INGENIEUR TECHNICO-COMMERCIAL H/F - Lyon (69)
0 Partages Entreprise reconnue dans la fourniture de machines-outils (tôlerie et mécanique), notre client se distingue par la technicité de ses solutions et la qualité de son accompagnement auprès…
Juriste urbanisme / immobilier / foncier (H/F)
Détails de l'offre Famille de métiers Urbanisme, aménagement et action foncière Domanialité e…
Ingénieur(e) d'affaires Senior F/H
POURQUOI REJOINDRE LEAP ? - Un projet entrepreneurial fort : Tu rejoindras une structure en pleine création où ton impact se fera sentir dès le premier jour. - Un environnement dynamique et agile …
Junior Strategy Consultant - Private Equity - Lyon F/H
Description Who we are: Indefi is a fast-growing strategy consulting firm that serves financial investors (private equity specialists, generalist asset managers, wealth managers, etc.). Establi…
Intégrateur DevSecOps (F/H)
A LIRE ATTENTIVEMENT AVANT DE POSTULER ⬇ 2 jours de télétravail / semaine - Lyon (03) - Expérience de 3 ans minimum Vous aimez les environnements de production où intégration, exploitation et …
Réceptionniste night H/F - CDI
Description de l'entreprise Prêt(e) à faire vibrer les nuits parisiennes avec nous ? Rejoins le Mercure Paris Gare de Lyon en devenant notre réceptionniste night et deviens l’ambassadeur(trice…
Data Analyst Senior - Lyon - Démarrage 1er Janvier 2027 - Mission Longue - Freelance
Taux journalier (TJM): 500 Dans le cadre d'un programme data d'envergure, notre client — un acteur de référence — souhaite intégrer l'IA au cœur de ses chantiers data : valorisation de documents et d…