Vacancy Description
Responsibilities
- Conduct evaluation research to turn eval targets into original benchmark designs
- Own measurement questions including construct validity, item discrimination, and contamination
- Build evaluation packages with expert-verified ground truth and rigorous QC
- Recruit, calibrate, and review a pool of subject‑matter experts across coding, agentic, and STEM fields
- Act as a technical point of contact for AI labs to translate measurement needs into evaluation designs
- Turn research findings into public benchmarks and papers for venues like NeurIPS, ICLR, and ACL
Requirements
- Research background in ML evaluation or benchmarking
- Deep expertise in LLM/frontier-model benchmarking, specifically code‑model and agentic evaluation
- Strong understanding of psychometrics, rubrics, and the measurement problem
- Interest in AI safety, capability elicitation, and robustness <...
Ready to Apply?
अभी आवेदन करें
Submit your application for Research Scientist at Remotedxb
Apply for this Position