Baseten Logo

Baseten

Site Reliability Engineer

Reposted 5 Days Ago
Remote or Hybrid
Hiring Remotely in Montréal, QC, CAN
Mid level
Remote or Hybrid
Hiring Remotely in Montréal, QC, CAN
Mid level
As an AI Support Engineer, you'll manage support requests, resolve user issues, optimize ML models, and contribute to product development.
The summary above was generated by AI

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products.

THE ROLE

As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently.

You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company.

EXAMPLE INITIATIVES

You'll work on projects like these as part of the SRE team:

  • Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services.

  • Building AI-assisted tooling for incident triage and response.

RESPONSIBILITIES

  • Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking.

  • Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code.

  • Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution.

  • Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations.

  • Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, concurrency, and model lifecycle management.

  • Define and instrument SLOs and SLIs across customer workloads and internal services.

  • Navigate ambiguity, make principled tradeoffs, and avoid unnecessary complexity in the systems you build and the processes you define.

REQUIREMENTS

  • Extensive hands-on experience with Kubernetes (multi-cloud experience across EKS, GKE, or similar is a strong plus).

  • Experience in building and maintaining scalable infrastructure.

  • Strong foundation in observability tooling: metrics (VictoriaMetrics, Prometheus), logging (Loki, ELK), dashboards (Grafana), and alerting pipelines. Observability-as-code experience is a plus.

  • Experience with infrastructure-as-code (Terraform, Helm) and GitOps workflows (Flux CD, ArgoCD).

  • Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis.

  • Comfort working at the intersection of engineering and operations — you write code, but you also think deeply about process, escalation paths, and operational leverage.

  • Familiarity with incident management platforms (incident.io or similar) is a plus.

  • No prior ML experience required, but curiosity about how ML models are deployed and served at scale will serve you well.

BENEFITS

  • Competitive compensation, including meaningful equity.

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

  • Paid parental leave

  • Fertility and family-building stipend through Carrot

  • Company-facilitated 401(k)

  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Similar Jobs

7 Days Ago
In-Office or Remote
Canada
Senior level
Senior level
Big Data • Information Technology • Software • Analytics • Energy
Manage and scale Enverus' global AWS infrastructure, automate deployments and CI/CD, ensure high uptime, collaborate with developers to enable zero-downtime releases, participate in on-call rotations, and improve operational practices.
Top Skills: AWSAzureC#Ci/CdCloudFormationGoKubernetesLinuxPythonTerraformWindows
12 Days Ago
Easy Apply
Remote
Canada
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Maintain and improve reliability, scalability, and automation for user-facing production systems. Build infrastructure tooling, operate Kubernetes-based services, write IaC, participate in on-call and incident response, and advance observability and runbooks to reduce toil and improve platform reliability.
Top Skills: AWSCi/CdGCPGitopsGoInfrastructure As Code (Iac)KubernetesKubernetes Operators/ControllersLoggingMetricsRubySlos/SlisTerraform
8 Days Ago
Easy Apply
Remote or Hybrid
Easy Apply
Senior level
Senior level
Artificial Intelligence • Marketing Tech • Software
Hands-on Senior Site Reliability Engineer responsible for improving infrastructure tooling and automation, building and supporting core applications, monitoring capacity and performance, troubleshooting incidents, participating in on-call rotations, and collaborating across teams to scale a multi-cloud, multi-region content serving platform and advance observability and SLO frameworks.
Top Skills: AWSEksGCPGkeGoGrafana AlloyKubernetesLinuxLokiNode.jsPrometheusPythonRubyShell ScriptingTempoTerraformThanos

What you need to know about the Montreal Tech Scene

With roots dating back to 1642, Montreal is often recognized for its French-inspired architecture and cobblestone streets lined with traditional shops and cafés. But what truly sets the city apart is how it blends its rich tradition with a modern edge, reflected in its evolving skyline and fast-growing tech industry. According to economic promotion agency Montréal International, the city ranks among the top in North America to invest in artificial intelligence, making it le spot idéal for job seekers who want the best of both worlds.

Key Facts About Montreal Tech

  • Number of Tech Workers: 255,000+ (2024, Tourisme Montréal)
  • Major Tech Employers: SAP, Google, Microsoft, Cisco
  • Key Industries: Artificial intelligence, machine learning, cybersecurity, cloud computing, web development
  • Funding Landscape: $1.47 billion in venture capital funding in 2024 (BetaKit)
  • Notable Investors: CIBC Innovation Banking, BDC Capital, Investissement Québec, Fonds de solidarité FTQ
  • Research Centers and Universities: McGill University, Université de Montréal, Concordia University, Mila Quebec, ÉTS Montréal

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account