Veröffentlicht am 20. Sept. 2026 · Wir haben am 21. Sept. 2026 bestätigt, dass er noch aktiv ist
Gehört dieses Unternehmen Ihnen?US$ 15 – US$ 30 pro Projekt
We are seeking experienced AI Evaluators, QA Specialists, and Technical Pod Leads to evaluate and benchmark frontier LLMs and agentic AI systems for client pilot projects at JudgeMyAI. Key Responsibilities: • Evaluate AI model outputs across multi-turn reasoning, agentic tool-use, and code generation (SWE-bench). • Stress-test models via adversarial prompts, edge-case analysis, and red-teaming. • Apply and calibrate evaluation rubrics (RLHF, SFT, accuracy, safety, and hallucination checks). • Review and annotate dataset batches with high inter-annotator agreement.
• Experience with LLM evaluation, prompt engineering, or AI benchmarking (prior work on platforms like Turing, Outlier, micro1 or enterprise AI labs is a strong plus). • Background in Python, Software Engineering, Data Science, or Advanced STEM. • Strong analytical reasoning and strict attention to detail. Engagement Details: • Remote and flexible hours (asynchronous execution). • Fixed milestone or hourly payouts in USD.
Erstellen Sie ein kostenloses Konto, um die vollständige Stelle zu sehen und sich zu bewerben.