Vacancy Description
We are sharing a specialised full-time consulting opportunity for experienced QA and test engineers with strong expertise in test-case design, end-to-end debugging, quality assurance, Python, and complex technical evaluation workflows.
This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will review complex multi-step tasks, test reference solutions, identify ambiguity and grading gaps, debug technical environments, and develop repeatable quality processes that keep benchmark results accurate and trustworthy.
Key Responsibilities
Test Case Design
- Create comprehensive test cases confirming that benchmark tasks function as intended
- Design positive, negative, boundary, and edge-case tests
- Validate task requirements, expected outputs, reference solutions, and grading logic
- Identify scenarios that may produce in...