FirstPrinciples Foundation Logo

FirstPrinciples Foundation

AI & HPC Infrastructure Engineer

Reposted 20 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in Ontario, ON, CAN
Senior level
Remote
Hiring Remotely in Ontario, ON, CAN
Senior level
Design, deploy, and operate Kubernetes, GPU, and HPC compute across bare metal and cloud. Manage scheduling (Slurm/K8s), CUDA and drivers, storage, observability, IaC, deployment workflows, and support researchers to ensure reliable, scalable inference and research infrastructure.
The summary above was generated by AI
About FirstPrinciples

FirstPrinciples is a research organization building AI infrastructure for discovery in fundamental science. Currently, our work focuses on building systems like Theo, the AI Physicist, which is a domain-specialized system for research in fundamental physics.

We’re a fast-growing, remote-first team of builders, researchers, engineers, and thinkers working across Canada, the US, the UK, and expanding globally. What brings us together is a shared curiosity about how the universe works, and a belief that we can build systems that help us explore it more effectively.

We spend our time working on questions that don’t have clear answers, like how to design AI that can reason through scientific problems, and how the scientific process as a whole might evolve. This is work that sits somewhere between creativity and rigorous thinking, and often requires comfort with ambiguity and iteration. If you’re someone who enjoys tackling big, abstract problems and building the infrastructure that makes ambitious research possible, you’ll likely find the work here interesting.

Why This Role Exists

We’re building the next generation of infrastructure for AI-driven scientific discovery, and we need someone who can help own the systems that make our research and inference workloads reliable, scalable, and fast.

This role is about building and operating the compute foundation behind our AI Physicist: Kubernetes clusters, Linux systems, GPU infrastructure, cloud environments, HPC-style compute, deployment workflows, monitoring, and automation. As our workloads grow, we need infrastructure that can support both experimentation and production-like inference across cloud, bare metal, and hybrid environments.

You’ll play a central role in shaping how we run compute at FirstPrinciples. That includes provisioning and managing clusters, improving reliability and observability, reducing operational toil, supporting researchers and engineers, and helping us make practical decisions about when to use managed cloud services, self-managed Kubernetes, Slurm-style systems, or owned hardware.

We’re looking for someone hands-on, systems-oriented, and comfortable working in a fast-moving research environment. You should have strong Kubernetes and Linux fundamentals, good operational instincts, and enough experience with cloud and HPC/GPU infrastructure to help us build toward a robust bare metal and multi-cloud inference platform.

What You’ll Do
  • Design, deploy, and operate Kubernetes infrastructure for AI inference, research, and engineering workloads

  • Set up and manage GPU and HPC-style compute environments, including scheduling, utilization, job management, and node-level troubleshooting

  • Work with systems such as Kubernetes, Slurm or similar schedulers, container runtimes, GPU drivers & libraries (ie; CUDA), storage systems, and observability tools

  • Build and manage Linux-based compute environments, including provisioning, networking, storage, monitoring, access control, and lifecycle management

  • Help architect bare metal, cloud, and hybrid infrastructure across AWS, GCP, Azure, or equivalent platforms

  • Own the reliability and operational health of infrastructure systems, including monitoring, alerting, incident response, capacity planning, and performance tuning

  • Improve deployment workflows, automation, configuration management, secrets management, and infrastructure-as-code practices

  • Partner with ML engineers, researchers, and software engineers to understand workload requirements and translate them into practical infrastructure designs

  • Evaluate tradeoffs between managed cloud services, self-managed Kubernetes, HPC schedulers, bare metal deployments, and multi-cloud architectures

  • Build tooling, documentation, runbooks, and operational practices that help the team move quickly without making infrastructure fragile or opaque

  • Balance speed and robustness, knowing when to prototype quickly and when to harden systems for long-term use

Who You Are
  • Strong infrastructure builder with experience operating production, research, cloud, or high-performance compute systems

  • Deeply comfortable with Linux administration, including debugging networking, storage, system services, permissions, performance issues, and node-level failures

  • Experienced with Kubernetes in real environments, including cluster operations, deployments, networking, observability, scaling, and troubleshooting

  • Comfortable working with cloud infrastructure on AWS, GCP, Azure, or equivalent platforms

  • Familiar with infrastructure automation and configuration tools such as Terraform, Ansible, Helm, ArgoCD, GitOps workflows, or similar systems

  • Experienced with GPU-heavy, compute-heavy, or HPC-style workloads, especially in environments involving AI, ML, research computing, or scientific workloads

  • Able to work across bare metal and cloud environments, and interested in the practical tradeoffs between the two

  • Comfortable reasoning about resource scheduling, cluster utilization, autoscaling, storage, networking, and observability for distributed workloads

  • Practical and ownership-oriented; you can take ambiguous infrastructure needs and turn them into working systems

  • Comfortable collaborating across disciplines, especially with researchers and engineers who may not think in infrastructure terms

  • Able to operate independently as a senior or strong intermediate contributor, while knowing when to bring others into important technical decisions

  • Motivated by building foundational systems that make ambitious technical and scientific work possible

Bonus
  • Hands-on experience with production-grade LLM inference and serving engines, such as vLLM, SGLang, or TensorRT

  • Experience working at an AI company, ML infrastructure team, research lab, university compute environment, HPC center, or scientific computing organization

  • Experience supporting model inference, model serving, distributed training, high-throughput batch workloads, or internal ML platforms

  • Hands-on experience with Slurm or similar HPC schedulers, including job scheduling, resource allocation, queue management, and cluster configuration

  • Experience operating GPU infrastructure, including NVIDIA drivers, CUDA, container runtimes, scheduling, utilization, and hardware failure modes

  • Experience with RDMA, InfiniBand, high-performance networking, distributed filesystems (ie. Lustre, BeeGFS), object storage, or storage systems for compute-heavy workloads

  • Experience with Kubernetes operators, custom controllers, CRDs, or platform tooling for AI/ML workloads

  • Experience with Prometheus, Grafana, Loki, OpenTelemetry, Datadog, or similar monitoring, logging, and observability tools

  • Experience with container registries, image optimization, CI/CD systems, deployment pipelines, and secure software delivery

  • Experience leading engineering operations or infrastructure efforts while remaining hands-on technically

  • Familiarity with security, access control, secrets management, and reliability practices in production or research environments

What You’ll Get
  • The opportunity to work on foundational problems at the intersection of AI and physics

  • A high-trust, low-bureaucracy environment with real ownership

  • Remote-first work with flexibility in how you structure your day

  • Exposure to cutting-edge ideas across AI, scientific discovery, infrastructure, and emerging technologies

  • A culture that values curiosity, depth of thinking, and first-principles reasoning

  • The chance to shape the compute and inference infrastructure behind advanced AI systems for scientific discovery

Similar Jobs

4 Minutes Ago
Remote or Hybrid
CA
Senior level
Senior level
Gaming
Lead UX/UI design for mobile games: create flows, wireframes, Figma prototypes, design systems, visual assets, and Unity implementations. Conduct research and testing, collaborate with art and engineering, mentor UX staff, and own UX across multiple feature pods to deliver engaging, usable game experiences.
Top Skills: Adobe Creative SuiteFigmaUnity
4 Minutes Ago
Remote or Hybrid
CA
Senior level
Senior level
Gaming
Lead design and implementation of large-scale game systems for a new mobile title. Write UX stories and design docs, prototype gameplay, create balancing models and tools, collaborate across disciplines, analyze live metrics to iterate systems, and evangelize game design best practices.
23 Minutes Ago
Remote or Hybrid
CA
Senior level
Senior level
Blockchain • Fintech • Mobile • Payments • Software • Financial Services
As a Staff Software Engineer, you will lead technical strategies, drive architectural decisions, mentor engineers, and develop consumer banking solutions using various technologies.
Top Skills: AuroraAWSBuildkiteDatadogDynamoDBGradleGrpcGuiceHibernateHTTPJavaJettyJSONJunitKafkaKotlinMySQLOkhttpProtocol BuffersRedis

What you need to know about the Montreal Tech Scene

With roots dating back to 1642, Montreal is often recognized for its French-inspired architecture and cobblestone streets lined with traditional shops and cafés. But what truly sets the city apart is how it blends its rich tradition with a modern edge, reflected in its evolving skyline and fast-growing tech industry. According to economic promotion agency Montréal International, the city ranks among the top in North America to invest in artificial intelligence, making it le spot idéal for job seekers who want the best of both worlds.

Key Facts About Montreal Tech

  • Number of Tech Workers: 255,000+ (2024, Tourisme Montréal)
  • Major Tech Employers: SAP, Google, Microsoft, Cisco
  • Key Industries: Artificial intelligence, machine learning, cybersecurity, cloud computing, web development
  • Funding Landscape: $1.47 billion in venture capital funding in 2024 (BetaKit)
  • Notable Investors: CIBC Innovation Banking, BDC Capital, Investissement Québec, Fonds de solidarité FTQ
  • Research Centers and Universities: McGill University, Université de Montréal, Concordia University, Mila Quebec, ÉTS Montréal

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account