Kuailu Software (Singapore) Pte. Ltd. · MyCareersFuture · 6d
Large Language Model (LLM) Evaluation Engineer
Singapore, Central- Posted
- 2026-10-02 (6d)
- Place
- Singapore, Central
- Commitment
- Full Time
- Salary
- SGD 6,000 – 12,000 / month
- Experience
- 3+ YOE
- Education
- Bachelor's
- Department
- Engineering
- Source
- MyCareersFuture (the employer’s own listing)
Your match
Sign in to see how your skills match this job.
Skills in this posting
PythonRegression TestingCI/CDArtificial IntelligenceData AnalysisvLLM
Job Responsibilities
• Build and maintain an automated LLM evaluation pipeline covering multiple dimensions, including general capabilities, Agent capabilities, and persona/role-playing. The pipeline should support one-click evaluation, historical result comparison, and regression testing.
• Conduct general capability evaluations using benchmarks such as MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, including benchmark deployment, execution, and results analysis.
• Conduct Agent capability evaluations, including setting up evaluation environments and tracking metrics for benchmarks such as BFCL, τ-bench, and GAIA.
• Design and execute persona/role-playing evaluation frameworks, covering metrics such as identity recognition, role compatibility, multi-turn stability, and style consistency.
• Record and analyze evaluation results from training runs, conduct comparative analysis and anomaly detection, and produce checkpoint evaluation reports.
• Conduct regular intermediate evaluations during the pre-training stage to track the evolution and improvement of model capabilities.
Job Requirements
• Bachelor's degree or above in Computer Science, Artificial Intelligence, or a related field.
• Familiarity with mainstream LLM evaluation benchmarks and frameworks, such as lm-eval-harness, OpenCompass, and EvalPlus.
• Strong proficiency in Python, with the ability to independently build evaluation pipelines covering model inference/deployment, batch evaluation, and results analysis.
• Familiarity with LLM inference frameworks such as vLLM and SGLang, with the ability to deploy models for batch inference and evaluation.
• Experience in evaluation data analysis and visualization.
• Detail-oriented and rigorous, with a strong focus on ensuring the reproducibility and reliability of evaluation results.
Preferred Qualifications
• Experience with Agent evaluation, particularly BFCL, τ-bench, GAIA, or SWE-bench.
• Experience with persona or role-playing evaluation, such as CharacterBench or RMTBench.
• Experience with evaluation automation and CI/CD integration.
• Understanding of model training workflows, with the ability to understand the relationship between training checkpoints and evaluation results.
Also posted at mycareersfuture.gov.sg
Kuailu Software (Singapore) Pte. Ltd.
6 open roles in Singapore, straight from Kuailu Software (Singapore) Pte. Ltd.’s own careers page.
- Foundation Model Pre-training ScientistSingapore, Central · 6d
- Global Head of AI Compute InfrastructureSingapore, Central · 6d
- Talent Acquisition ExecutiveSingapore, Central · 6d
- GPU Data Center Sales / Business DevelopmentSingapore, Central · 2w
- AI Chief ScientistSingapore, Central · 2w