Boson AI Logo

Boson AI

Network Engineer, AI/ML Infrastructure

Posted 11 Days Ago
Be an Early Applicant
In-Office
Toronto, ON
Mid level
In-Office
Toronto, ON
Mid level
The Network Engineer will design, build, and optimize networking infrastructure for AI/ML operations, manage network fabrics, troubleshoot issues, and plan for capacity.
The summary above was generated by AI
About The Role

We're seeking an experienced Network Engineer to design, build, and optimize the high-performance networking infrastructure powering our AI/ML operations in Toronto. You'll work at the cutting edge of network technology—managing InfiniBand and ultra-high-speed Ethernet fabrics that connect NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage, and hundreds of servers.

You'll be hands-on with the full lifecycle of our network infrastructure: planning, building, testing, deploying, and keeping everything running at peak performance. That means troubleshooting issues as they arise, monitoring network performance and throughput, developing automation to streamline operations, and working closely with HPC and ML teams to ensure they have the bandwidth they need. You'll also help us plan for future capacity and evaluate emerging network technologies as we scale to meet increasingly demanding workloads.

Responsibilities

  • Configure and maintain InfiniBand and high-speed Ethernet fabrics
  • Optimize network performance for RDMA, and GPU-to-GPU communication
  • Manage network switches (Mellanox, NVIDIA, Micas Networks)
  • Troubleshoot network bottlenecks and latency issues
  • Plan and execute network upgrades and expansions
  • Network security implementation (firewalls, VLANs, ACLs)
  • Collaborate on storage network optimizationInfrastructure monitoring

Minimum Qualifications

  • 4+ years of network engineering experience in production environments
  • Strong understanding of L2/L3 networking protocols (TCP/IP, BGP, OSPF, VLANs)
  • Hands-on experience with high-speed networking (100Gb+ Ethernet and InfiniBand)
  • Hands-on experience with network security (firewalls, ACLs, network segmentation)
  • Knowledge of HPC network topologies
  • Experience with InfiniBand fabrics including RDMA, RoCE, IPoIB
  • Strong troubleshooting and problem-solving skills

Preferred Qualifications

  • Experience in data center environments or AI/ML infrastructure
  • Hands-on experience with high-performance Ethernet switches (e.g., Broadcom Tomahawk), and latest InfiniBand switches (e.g., Nvidia/Mellanox)
  • Experience optimizing networks for GPU-to-GPU communication
  • Experience with open-source firewall solutions (OPNsense, pfSense, or similar)
  • Experience with network automation tools
  • Understanding of distributed storage networking (Ceph cluster networks)
  • Familiarity with network monitoring and observability tools (Prometheus, Grafana)
  • Knowledge of multi-site network connectivity and WAN optimization
  • Familiarity with cloud networking in at least one platform (AWS, GCP, or Azure) including VPC design, site-to-site VPN configuration, Direct Connect/ExpressRoute/Cloud Interconnect, hybrid cloud connectivity, and cloud-to-datacenter network integration

If you're a natural problem-solver with a passion for continuous learning, we'd love to hear from you.

Top Skills

AWS
Azure
Bgp
Broadcom Tomahawk
Ceph
Cloud Networking
Ethernet
GCP
Grafana
Infiniband
Ipoib
Mellanox
Nvidia
Opnsense
Ospf
Pfsense
Prometheus
Rdma
Roce
Tcp/Ip
Vlans

Similar Jobs

An Hour Ago
Easy Apply
Hybrid
2 Locations
Easy Apply
Mid level
Mid level
Big Data • Cloud • Software • Database
The role involves enhancing developer productivity by creating solutions for cross-team challenges, building tools, and providing support within MongoDB's engineering environment. It requires collaboration, understanding customer needs, and mentoring junior team members.
Top Skills: AWSC++GoKubernetesPython
An Hour Ago
Easy Apply
Hybrid
Toronto, ON, CAN
Easy Apply
Junior
Junior
Big Data • Cloud • Software • Database
The Legal Operations Analyst will manage CLM tools, drive process improvements, support legal teams, and promote best practices while collaborating cross-functionally.
Top Skills: BrightflagGoogle SuiteMalbekMS OfficeSalesforce
3 Hours Ago
In-Office
Toronto, ON, CAN
Senior level
Senior level
Cloud • Fintech • Food • Information Technology • Software • Hospitality
The Counsel, International Product will advise on product and marketing laws, ensure regulatory compliance, and facilitate international expansion for Toast's products.

What you need to know about the Singapore Tech Scene

The digital revolution has driven a constant demand for tech professionals across industries like software development, data analytics and cybersecurity. In Singapore, one of the largest cities in Southeast Asia, the demand for tech talent is so high that the government continues to invest millions into programs designed to develop a talent pipeline directly from universities while also scaling efforts in pre-employment training and mid-career upskilling to expand and elevate its workforce.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account