OKX Logo

OKX

Big Data Engineer, Web3

Reposted Yesterday
Be an Early Applicant
In-Office
Singapore, SGP
Mid level
In-Office
Singapore, SGP
Mid level
Operate and support multi-cloud data platform components (Alibaba Cloud and AWS). Build monitoring, SLA/SLOs, respond to incidents, troubleshoot and root-cause issues, optimize compute/storage and cloud costs, develop automation and operational tooling, and collaborate with data and platform teams and cloud vendors to improve platform reliability.
The summary above was generated by AI
OKX will be prioritising applicants who have a current right to work in Singapore, and do not require OKX's sponsorship of a visa.
Who We Are
At OKX, we believe that the future will be reshaped by crypto, and ultimately contribute to every individual's freedom.
 
OKX is a leading crypto exchange, and the developer of OKX Wallet, giving millions access to crypto trading and decentralized crypto applications (dApps). OKX is also a trusted brand by hundreds of large institutions seeking access to crypto markets. We are safe and reliable, backed by our Proof of Reserves. 
 
Across our multiple offices globally, we are united by our core principles: We Before Me, Do the Right Thing, and Get Things Done. These shared values drive our culture, shape our processes, and foster a friendly, rewarding, and diverse environment for every OK-er.
OKX is part of OKG, a group that brings the value of Blockchain to users around the world, through our leading products OKX, OKX Wallet, OKLink and more.
 
About The Opportunity
We are building the foundational data infrastructure that powers one of the world's leading crypto exchanges. As a Big Data Platform Engineer, you will design, build, and evolve the core platform services that enable hundreds of data engineers and analysts to move fast and ship reliable pipelines. You will be at the forefront of our AI agent initiative — embedding LLM-driven capabilities directly into the platform layer so that scheduling, cost optimization, and incident response increasingly run autonomously.
What You’ll Be Doing 
  • Platform Core: Design and operate large-scale distributed data systems
  • Own the big data compute and storage infrastructure (MaxCompute/ODPS, Hologres, Spark)
  • Build and maintain multi-site task orchestration that dynamically selects engines and enforces policy
  • Drive reliability and performance improvements across batch and real-time pipelines

  • AI Integration: Build the AI-native platform layer
  • Develop and expose MCP (Model Context Protocol) tool interfaces so AI agents can interact with platform APIs
  • Build the scheduling and cost-optimization agents that auto-tune resource allocation and alert severity
  • Instrument platform telemetry to feed AI-driven SLA monitoring and anomaly detection
  • Design context retrieval pipelines (RAG / vector search) for SQL code and config knowledge bases

  • Tooling & DX: Evolve the developer experience
  • Own the internal data development platform — IDE integrations, code review automation, deployment tooling
  • Build APIs-first tools (backfill, ingestion automation) designed for future MCP integration
  • Collaborate with data warehouse and service teams to define platform contracts

  • Ops & Governance: Drive operational excellence
  • Establish SLA benchmarks, cost metrics, and latency dashboards as AI optimization targets
  • Build automated incident response and root-cause analysis pipelines
  • Define and enforce infrastructure policies across multi-cloud environments
AI Agent Ownership — Platform Tier
  • Scheduling Agent: auto-configure task dependencies, engine selection, cost/performance trade-offs, and alert tiers
  • Operations Agent: detect pipeline latency, performance degradation, and schema drift; trigger remediation
  • Incident Response Agent: trace SLA breaches to root cause, assign accountability, generate post-mortems
  • MCP Tool Layer: design and maintain the cross-platform tool interfaces that all agents call into
What We Look For In You 
  • 5+ years of experience building large-scale data platforms (Hadoop/Spark/Flink or equivalent)
  • Deep expertise in distributed storage and compute systems (MaxCompute, Hologres, ClickHouse, Hive)
  • Strong software engineering skills in Java, Scala, or Python; experience with API-first design
  • Hands-on experience with task scheduling systems (Airflow, DolphinScheduler, or in-house equivalents)
  • Solid understanding of multi-cloud architectures and cost governance
  • Familiarity with LLM integration patterns: tool calling, RAG pipelines, context management
  • Experience with MCP or similar agent-tool frameworks is a strong plus
  • Passion for building systems that make other engineers 10x more productive
Perks & Benefits 
  • Competitive total compensation package
  • L&D programs and education subsidy for employees' growth and development
  • Various team building programs and company events
  • Wellness and meal allowances
  • Comprehensive healthcare schemes for employees and dependants
  • More that we love to tell you along the process!

Notice:
All official OKX vacancies are published on this website. While roles may appear on selected third-party platforms from time to time, information on other sites may be inaccurate or outdated. If in doubt, please apply directly through our official careers website.
Information collected and processed as part of the recruitment process of any job application you choose to submit is subject to OKX's Candidate Privacy Notice.
HQ

OKX Singapore, Singapore, SGP Office

Singapore, Singapore

Similar Jobs

Yesterday
In-Office or Remote
Singapore, SGP
Senior level
Senior level
Fintech • Financial Services • Cryptocurrency
Design and own large-scale Web3 big data platform and pipelines (real-time and batch), maintain data warehouse/lake, integrate AI/LLM capabilities (RAG, embeddings, Text2SQL), collaborate with AI, product, risk and growth teams, and optimize performance, reliability, and data quality for on-chain analytics and ML use cases.
Top Skills: Ai AgentsClickhouseDorisDuneEmbedding PipelinesEtherscan ApiFeature PlatformsFlinkHadoopHbaseHiveJavaKafkaLlmsMcpMilvusNode RpcsPgvectorPineconePythonRagReal-Time Feature ServingScalaSparkSQLText2SqlThe GraphVector DatabasesWeaviate
7 Hours Ago
In-Office
Singapore, SGP
Junior
Junior
Artificial Intelligence • Hardware • Information Technology • Machine Learning
The Staff Engineer will innovate and optimize 3D NAND equipment, oversee installations, manage equipment projects, and ensure compliance while collaborating with cross-functional teams.
Top Skills: 3D NandDocumentation SystemsProcess DevelopmentProject ManagementSemiconductor EquipmentTooling
7 Hours Ago
In-Office
Singapore, SGP
Senior level
Senior level
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Lead global Diffusion HVM process engineering to improve manufacturing capability, productivity, yield, and equipment performance. Drive supplier engagement, equipment technology development, cross-functional alignment, and mentor a team to meet technology and cost leadership goals.
Top Skills: Diffusion ProcessDramEquipment RoadmapsHigh-Volume Manufacturing (Hvm)NandProcess Control

What you need to know about the Singapore Tech Scene

The digital revolution has driven a constant demand for tech professionals across industries like software development, data analytics and cybersecurity. In Singapore, one of the largest cities in Southeast Asia, the demand for tech talent is so high that the government continues to invest millions into programs designed to develop a talent pipeline directly from universities while also scaling efforts in pre-employment training and mid-career upskilling to expand and elevate its workforce.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account