Veeda AI Logo

Veeda AI

Senior AI Infrastructure Engineer — HPC & Compute Clusters

Posted 2 Days Ago
Remote or Hybrid
Hiring Remotely in Zürich
Senior level
Remote or Hybrid
Hiring Remotely in Zürich
Senior level
Design, deploy, and maintain large-scale bare-metal GPU clusters for distributed training. Tune Slurm schedulers and Kubernetes integrations, optimize storage and high-speed networking, and build observability pipelines to ensure reliable, high-throughput GPU workloads for distributed deep learning.
The summary above was generated by AI
About US

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities
  • GPU Cluster Orchestration: Design, deploy, and maintain large-scale bare-metal GPU clusters running Slurm and Kubernetes to power multi-node distributed training runs.

  • Workload Management & Scheduler Tuning: Configure, tune, and optimize Slurm schedulers, dynamic job queuing, priority topologies, and autoscaling for optimal GPU utilization and job throughput.

  • Kubernetes Integration: Manage K8s clusters and hybrid Slurm-K8s environments (e.g., KubeRay, MPI Operator, Slurm on K8s) for serving, data processing pipelines, and interactive research environments.

  • Storage & Network Performance: Optimize high-speed interconnects (InfiniBand/RoCE, NCCL, NVLink) and high-throughput distributed storage (Lustre, WEKA, Ceph) to keep tens of thousands of GPUs continuously saturated.

  • Reliability & Monitoring: Build observability pipelines (Prometheus, Grafana, DCGM) to proactively detect hardware degradation, GPU silent errors, network flaps, and node failures before they impact training jobs.

Requirements
  • You have a Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in high-performance computing (HPC) or infrastructure engineering.

  • You have deep hands-on expertise administering Linux-based HPC clusters running Slurm or Kubernetes (job scheduling, cgroups, fair-share policies, topology configuration).

  • You possess strong troubleshooting skills in low-level Linux networking, kernel parameters, hardware diagnostics, and storage systems (e.g., NFS, NVMe-oF, or distributed filesystems).

  • You are proficient in automation and infrastructure-as-code tools (e.g., Ansible, Terraform, Helm, Python/Bash scripting).

  • Background supporting distributed deep learning frameworks (DeepSpeed, Ray).

Nice to Have
  • Experience managing high-density GPU infrastructure (NVIDIA H100/B200/B300 systems, DGX/HGX architectures, DCGM, NCCL tuning).

  • Experience with high-speed network fabrics including InfiniBand (Subnet Manager, SMDB) and RoCE (v2).

Similar Jobs

5 Hours Ago
Remote
Canada
Senior level
Senior level
Artificial Intelligence • Productivity • Software • Automation
As a Sr. Applied AI Engineer at Zapier, you will build and enhance AI platform capabilities, focusing on LLM Ops and ML Ops to support scalable AI development across teams.
Top Skills: Cloud InfrastructureLlm OpsMl OpsPythonTypescript
5 Hours Ago
Remote
Expert/Leader
Expert/Leader
Information Technology
As the Director of Data Science, you'll mentor teams, oversee projects, drive data-driven decisions, and implement AI and ML solutions while enhancing our data pipeline.
Top Skills: BigQueryClickhouseDruidMlNlpPythonRedshiftSQL
6 Hours Ago
Easy Apply
Remote
Easy Apply
Senior level
Senior level
Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Serve as Finance's single DRI for product launches with financial impact: evaluate new trade strategies, document fund flows and accounting treatments, design and implement financial controls, coordinate cross-functional stakeholders (accounting, tax, treasury, legal, operations), and partner with IT/vendors to automate and scale workflows for audit and regulatory readiness.
Top Skills: Generative Ai

What you need to know about the Montreal Tech Scene

With roots dating back to 1642, Montreal is often recognized for its French-inspired architecture and cobblestone streets lined with traditional shops and cafés. But what truly sets the city apart is how it blends its rich tradition with a modern edge, reflected in its evolving skyline and fast-growing tech industry. According to economic promotion agency Montréal International, the city ranks among the top in North America to invest in artificial intelligence, making it le spot idéal for job seekers who want the best of both worlds.

Key Facts About Montreal Tech

  • Number of Tech Workers: 255,000+ (2024, Tourisme Montréal)
  • Major Tech Employers: SAP, Google, Microsoft, Cisco
  • Key Industries: Artificial intelligence, machine learning, cybersecurity, cloud computing, web development
  • Funding Landscape: $1.47 billion in venture capital funding in 2024 (BetaKit)
  • Notable Investors: CIBC Innovation Banking, BDC Capital, Investissement Québec, Fonds de solidarité FTQ
  • Research Centers and Universities: McGill University, Université de Montréal, Concordia University, Mila Quebec, ÉTS Montréal

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account