Nanyang Technological University Logo

Nanyang Technological University

AI Engineer, Platforms

Reposted 16 Days Ago
Be an Early Applicant
In-Office
Singapore, SGP
Junior
In-Office
Singapore, SGP
Junior
Operate and optimize multi-cloud GPU inference platform for LLMs, build API services and CI/CD, manage AI clusters with IaC, improve observability, automate operations and build AI-assisted internal tools.
The summary above was generated by AI

AI Singapore (AISG) is a national AI programme launched by the National Research Foundation (NRF), Singapore, to build and anchor deep national capabilities in AI. AISG is supported through a government-wide partnership including the NRF, Ministry of Digital Development and Information (MDDI), Infocomm Media Development Authority (IMDA), Economic Development Board (EDB) and Enterprise Singapore (ESG). We bring together research institutions and the vibrant ecosystem of AI start-ups and companies to support impactful research, develop talent, and power Singapore's AI efforts.

This position will be hosted at the Nanyang Technological University (NTU) under VP (Artificial Intelligence & Digital Economy)’s office and we welcome you to join our community.

We're looking for an AI Engineer to join the Platform team within AI Products at AISG. In this role, you will be collaborating with different internal teams to design and implement optimized inference workflows and support the team to build customized LLM-based solutions.

Your work will directly contribute to the deployment, optimisation and management of large language models (LLMs) in production, retrieval-augmented generation (RAG) services, AI agent orchestration platforms and GPU-enabled AI infrastructure.

Responsibilities:

Platform operations and reliability

  • Own day-to-day operations of SEA-LION API Farm, our multi-cloud LLM inference platform — monitoring environment health, GPU capacity, performance, cost, and security posture.

  • Optimise LLM inference across various modalities to drive business value and support production goals.

  • Diagnose and troubleshoot performance and reliability issues on API Farm.

  • Build new API services such as batch API services, MCP services.

Infrastructure, CI/CD, and automation

  • Manage high performance AI clusters and storage systems using infrastructure-as-code (e.g. Terraform) across different cloud providers.

  • Develop and maintain CI/CD pipelines, container build/registry workflows, and deployment automation so teams can ship safely and frequently.

  • Strengthen observability across the stack including logs, metrics, traces, and dashboards, and reduce toil by automating repetitive operational tasks.

AI-assisted ops and continuous improvement

  • Use AI tools (e.g. Claude, Copilot, Cursor) appropriately in your daily work responsibilities.

  • Build internal tools leveraging AI to reduce manual effort in day-to-day operations.

Requirements:

You should be a hands-on engineer who is comfortable operating cloud and GPU infrastructure end-to-end, who understands how to deploy and run large language models reliably in production, and who actively uses AI tools to make platform work faster and more reliable.

  • A degree in Computer Science, Information Technology, or equivalent.

  • 1–3 years of DevOps, SRE, or platform engineering experience, with a track record of operating production systems at scale.

  • Hands-on experience operating workloads on different cloud providers including IaC (e.g. Terraform), containers and orchestration (e.g. Docker, Kubernetes), and managed services for compute, storage, and networking.

  • Strong knowledge on Inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM, Transformers).

  • Hands-on experience deploying and serving LLMs in production — model serving, GPU scheduling, autoscaling, latency/throughput optimisation, and inference cost management.

  • REST API design, model context protocol (MCP), Internet authentication patterns (e.g. OAuth).

  • Strong fundamentals in CI/CD, observability (logs/metrics/traces), and incident response.

  • Demonstrated use of AI tools (e.g. Claude, Copilot, Cursor) in your day-to-day engineering — for code generation, review, debugging, and documentation — with a clear sense of where they help and where they don't.

  • Solid scripting/programming skills (e.g. Python, Bash) and comfortable reading other people's code across the stack.

  • Strong communication skills with the ability to explain technical concepts.

Good to Have:

  • Experience with multimodal AI models (e.g. vision language models, audio language models).

  • C/C++/Rust/Go or other relevant programming languages.

  • Contributions to open-source AI/ML projects.

We regret that only shortlisted candidates will be notified.

Hiring Institution: NTU
HQ

Nanyang Technological University Singapore, Singapore, SGP Office

Singapore, Singapore

Nanyang Technological University Singapore Office

Singapore

Similar Jobs

23 Days Ago
In-Office or Remote
Singapore, SGP
Senior level
Senior level
News + Entertainment
Lead and operate enterprise Azure and Microsoft 365 platforms, manage hybrid identity (Entra ID/AD), deliver and govern AI solutions (Azure OpenAI, Azure ML, Copilot), provide L2/3 support, secure and govern cloud resources, manage on‑prem infrastructure, and drive vendor, project, and stakeholder engagement to scale AI initiatives.
Top Skills: Azure Ai ServicesAzure Ai StudioAzure DevopsAzure Machine LearningAzure Openai ServiceBicepConditional AccessCopilot ExtensibilityCopilot StudioDefenderEntra IdExchange OnlineGithub ActionsHybrid Active DirectoryIntuneMfaMicrosoft 365 CopilotMicrosoft Graph ApiMicrosoft TeamsOauth 2.0OnedriveOpenid Connect (Oidc)Power AutomatePower PlatformPowershellPythonRest ApisSAMLSharepoint OnlineTerraform
An Hour Ago
In-Office
Singapore, SGP
Expert/Leader
Expert/Leader
Fintech • Information Technology • Financial Services
Designs, builds, and operates secure AWS-native data pipelines and platforms for batch and event-driven workloads. Develops Airflow workflows using Python and SQL, implements tested data transformations, and integrates enterprise, security, and operational data sources. The role embeds data quality, lineage, observability, governance, encryption, access controls, and monitoring. It also contributes to CI/CD, infrastructure as code, documentation, code reviews, incident analysis, and cross-functional solutions supporting risk and regulatory needs.
Top Skills: Amazon CloudwatchAmazon OpensearchAmazon S3Apache AirflowAWSAws GlueAws IamAws LambdaAws Secrets ManagerAws Step FunctionsCi/CdDbtGitGreat ExpectationsInfrastructure As CodePythonSQLTerraform
An Hour Ago
Easy Apply
Hybrid
Singapore, SGP
Easy Apply
Senior level
Senior level
Big Data • Cloud • Software • Database
Drive MongoDB growth by prospecting enterprise technology leaders, building relationships, managing full sales cycles, developing territory plans, generating pipeline, and achieving revenue targets. The role focuses on the Vietnam market and requires strong Vietnamese and English skills, enterprise technology sales expertise, cold-calling ability, and success closing new customers.
Top Skills: AWSGCPAzureMongoDBOutreachSales NavigatorSalesforceSendosoTerretZoominfo

What you need to know about the Singapore Tech Scene

The digital revolution has driven a constant demand for tech professionals across industries like software development, data analytics and cybersecurity. In Singapore, one of the largest cities in Southeast Asia, the demand for tech talent is so high that the government continues to invest millions into programs designed to develop a talent pipeline directly from universities while also scaling efforts in pre-employment training and mid-career upskilling to expand and elevate its workforce.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account