Founded in 2013, Fusemachines is a global provider of enterprise AI products and services, on a mission to democratize AI. Leveraging proprietary AI Studio and AI Engines, the company helps drive the clients’ AI Enterprise Transformation, regardless of where they are in their Digital AI journeys. With offices in North America, Asia, and Latin America, Fusemachines provides a suite of enterprise AI offerings and specialty services that allow organizations of any size to implement and scale AI. Fusemachines serves companies in industries such as retail, manufacturing, and government.Fusemachines continues to actively pursue the mission of democratizing AI for the masses by providing high-quality AI education in underserved communities and helping organizations achieve their full potential with AI.Type: Full-Time, RemoteAbout the Role
We are seeking a Senior Software Engineer to build and operate high-performance backend services using Java. This role is ideal for an experienced backend engineer who understands distributed systems, testing, observability, and production reliability, with enough machine learning knowledge to integrate applications with models and feature stores.
You will work on systems with demanding availability, latency, throughput, and service-level requirements. You will collaborate with software, machine learning, data, infrastructure, and product teams to deliver reliable production solutions.
This is primarily a software engineering role. Deep machine learning research experience is not required.
ResponsibilitiesDesign, build, test, and maintain scalable backend services using Java.
Develop APIs, microservices, and event-driven components for high-volume production systems.
Apply established Java architecture patterns, coding standards, and engineering best practices.
Integrate backend applications with machine learning models, inference endpoints, and feature stores.
Design reliable model-calling workflows with appropriate timeouts, retries, validation, and fallback behavior.
Build highly available, low-latency services that meet defined SLAs and service-level objectives.
Monitor and optimize latency, throughput, availability, resource utilization, and error rates.
Implement resilience patterns such as caching, circuit breakers, rate limiting, and graceful degradation.
Develop unit, integration, contract, performance, and end-to-end tests.
Implement logging, metrics, dashboards, distributed tracing, and actionable alerts.
Use platforms such as Datadog or similar tools to monitor systems and investigate production issues.
Participate in incident response, root-cause analysis, and reliability improvements.
Use AI-assisted coding tools responsibly to support development, testing, documentation, and debugging.
Participate in architecture discussions, technical design reviews, and code reviews.
Communicate technical decisions, risks, dependencies, and tradeoffs clearly to technical and nontechnical stakeholders.
Promote strong software engineering practices across the team.
Strong professional experience developing production applications with Java.
Experience designing backend services, APIs, microservices, or distributed systems.
Strong understanding of Java design patterns, object-oriented programming, and software architecture.
Experience building systems with demanding availability, scalability, throughput, or latency requirements.
Experience with automated testing and continuous integration and delivery practices.
Experience with relational or non-relational databases.
Experience implementing production logging, monitoring, metrics, tracing, and alerting.
Familiarity with Datadog, OpenTelemetry, Grafana, Prometheus, New Relic, or comparable platforms.
Understanding of SLAs, service-level indicators, service-level objectives, and production reliability.
Familiarity with cloud platforms, containers, and modern deployment environments.
Basic understanding of machine learning concepts and how applications interact with deployed models.
Strong troubleshooting, problem-solving, and production support skills.
Excellent written and verbal communication skills.
Experience integrating applications with machine learning inference services or feature stores.
Experience with low-latency or real-time decisioning systems.
Experience with Kafka, message queues, streaming platforms, or event-driven architectures.
Experience with Kubernetes, Redis, distributed caching, or performance testing.
Familiarity with model versioning, feature freshness, prediction logging, and controlled model rollouts.
Experience in advertising technology, digital marketplaces, auction systems, recommendation systems, or personalization.
Experience using AI-assisted development tools such as GitHub Copilot, Claude Code, or Cursor.
Fusemachines is an Equal Opportunities Employer, committed to diversity and inclusion. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or any other characteristic protected by applicable federal, state, or local laws.



