NCS · SmartRecruiters · 1mo
#EG Senior / LLMOps Engineer
Singapore- Posted
- 2026-08-21 (1mo)
- Place
- Singapore
- Commitment
- Full Time
- Experience
- 4+ YOE
- Department
- Others
- Source
- SmartRecruiters (the employer’s own listing)
Your match
Sign in to see how your skills match this job.
Skills in this posting
PythonTypeScriptGoRedisOpenSearchAmazon Web ServicesGoogle Cloud PlatformDockerKubernetesHelmCI/CDGitHub Actions
NCS is a leading AI Tech Services company. With a 15,000-strong team across the Asia Pacific, NCS scales its platforms and capabilities to provide clients with greater agility and AI expertise across a range of Industries. Embracing a strong ecosystem of global partners, NCS transforms technology services delivery combining AI with digital resilience to drive real business impact. NCS is a subsidiary of the Singtel Group.
This role sits within NCS AI Central's (AIC) Forward Deployed Engineering (FDE) model — the combined capability that takes AI solutions from proof-of-concept through to hardened production systems. You will operate across both fast-moving FDE engagements (POC/POV, pilot deployments for strategic and lighthouse clients) and steady-state system development and maintenance work — bringing the same rigor and a reusable, asset-fed approach to both.
What will you do?
1. Model & Deployment Operations
• Own the deployment pipeline for LLM and agentic applications — versioning models, prompts, and configurations as code, with safe rollout and rollback paths.
• Build repeatable CI/CD pipelines for AI services (containerised, on Kubernetes) so new model or prompt versions ship without manual intervention.
• Support rapid, disposable environment spin-up for FDE POC/POV work, then harden the same pipeline into a production-grade deployment when an engagement scales.
2. Monitoring & Reliability
• Instrument LLM applications with observability for latency, error rates, output drift, and hallucination signals — not just infrastructure uptime.
• Define and track SLAs/SLOs for production AI services in partnership with AI Architects and Solution Architects.
• Set up alerting and runbooks so issues in live AI systems are caught and triaged before clients notice.
3. Cost & Performance Management
• Monitor token spend and inference cost per model/engagement; flag anomalies and right-sizing opportunities in partnership with AI FinOps.
• Tune routing between model tiers (frontier vs. smaller/fine-tuned models) for cost-performance balance at production scale.
4. FDE & Development/Maintenance Coverage
• During FDE engagements: stand up lightweight, reusable deployment scaffolding that lets AI Engineers iterate quickly on POC/POV without operational overhead.
• During system development & maintenance engagements: take ownership of steady-state production operations, patching, upgrades, and incident response for live AI systems.
• Contribute reusable deployment patterns back into the shared internal asset library so future engagements start from a hardened baseline, not from zero.
5. Collaboration
• Work closely with AI Engineers, Cloud Architects, and AI Site-facing leads to ensure a smooth handoff from POC to Scale to Operate.
• Mentor engineers on LLMOps practices and participate in PRR (Production Readiness Review) gate reviews.
Role Levels We Are Hiring For
We are hiring at two levels for this role. All responsibilities above apply to both; the distinction is in scope of ownership, years of experience, and seniority of judgement expected.
LLMOps Engineer
• 2–4 years of relevant experience. Executes deployment pipelines, monitoring setup, and cost tracking for individual engagements, under guidance from a Senior LLMOps Engineer or AI Architect.
• Builds and maintains CI/CD and observability for one or two engagements at a time; escalates novel production incidents to senior team members.
Senior LLMOps Engineer
• 5+ years of relevant experience, including prior ownership of production AI/ML systems end-to-end. Sets LLMOps standards and reusable deployment patterns across multiple engagements.
• Leads incident response for critical production issues, mentors junior LLMOps and AI Engineers, and engages directly with client technical stakeholders on production-readiness and reliability.
The ideal candidate should possess:
• 2+ years in DevOps/MLOps/platform engineering, with hands-on exposure to LLM or ML systems in production (not just prototypes); 5+ years with end-to-end ownership expected at Senior level.
• Hands-on with containers, Kubernetes, CI/CD (GitHub Actions/GitLab/Jenkins), and Infrastructure-as-Code (Terraform).
• Practical experience deploying and operating LLM applications (RAG/agentic systems) at scale, including model gateway/routing patterns.
• Strong scripting/programming ability (Python and/or Go), comfortable working across cloud platforms (AWS/Azure/GCP; GCC/HCC exposure a plus).
• Working knowledge of observability tooling (OpenTelemetry, Prometheus/Grafana, ELK/OpenSearch) applied to AI-specific signals (drift, hallucination rate, token cost).
• Comfortable operating in both fast-paced, ambiguous POC/POV settings and disciplined, SLA-driven production support environments.
• Working knowledge of the China AI model/tech stack (e.g., DeepSeek, Qwen, GLM, Kimi, MiniMax) — deployment patterns, licensing, and self-hosting requirements.
Preferred Qualifications
• Experience with model gateways/routers (LiteLLM, Bedrock, Vertex, Azure OpenAI) and vector databases (pgvector, Pinecone, Weaviate).
• Exposure to regulated or government environments (IM8/VAPT, PDPA) and multi-tenant/data-residency patterns.
• Familiarity with agent orchestration frameworks (LangGraph, Semantic Kernel) and prompt/version control tooling.
• Experience contributing to or maintaining a reusable internal platform/asset library.
• Hands-on deployment/self-hosting experience with Chinese open-weight models (DeepSeek, Qwen, GLM) via vLLM/TGI or similar inference runtimes.
Tech Stack (Illustrative)
• Languages: Python, Go/TypeScript
• Platform: Docker, Kubernetes, Helm, Argo CD, Terraform, Vault
• LLM Runtime: OpenAI/Azure OpenAI/Bedrock/Vertex; DeepSeek/Qwen/GLM (China stack); vLLM/TGI; model gateways/routers
• Observability: OpenTelemetry, Prometheus/Grafana, ELK/OpenSearch, cost meters per request/model
• Storage/Search: Postgres, Redis; pgvector/Pinecone/Weavi…
NCS
200 open roles in Singapore, straight from NCS’s own careers page.
- [Uni - Jan till Jun 2027] Business & Strategy Analyst Intern (AI and Smart Systems)Singapore · Today
- #EG Full Stack Software Engineer (Must have Java)Singapore · Today
- #EG UI/UX DesignerSingapore · 1d
- #EG Senior AI Data EngineerSingapore · 1d
- #EG Cyber Engineer (Crowdstrike)Singapore · 2d
- #EX Event & Robotics Experience TechnicianSingapore · 2d
- Lead Consultant (OT Security)Singapore · 2d
- Cybersecurity Engineer (PKI, Cryptography Experience)Singapore · 2d
- Network Engineer [CCNA]Singapore · 2d
- [Uni – Jan till Jun 2027] Business Development Intern (Healthcare)Singapore · 3d
- Network & Security EngineerSingapore · 3d
- Consultant, IT Security (Cyber Access Management)Singapore · 3d