Baseten Logo

Baseten

Cloud Platform Engineer

Reposted 27 Days Ago
Remote or Hybrid
Hiring Remotely in Montréal, QC, CAN
Mid level
Remote or Hybrid
Hiring Remotely in Montréal, QC, CAN
Mid level
As a Site Reliability Engineer, you'll build and maintain infrastructure for ML models, automate processes, and collaborate cross-functionally.
The summary above was generated by AI

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.

THE ROLE

As a Cloud Platform Engineer, you'll envision and build robust systems and processes that ensure our infrastructure is scalable, reliable, and efficient. This can range from automating deployments and monitoring systems to optimizing performance and managing incidents.

We all work closely with our users, learning from their past struggles in operationalizing ML, onboarding them onto our platform, and turning our learnings into ideas for improving Baseten.

EXAMPLE INITIATIVES

You'll get to work on these types of projects as part of our Infrastructure team:

  • Multi-cloud capacity management

  • Inference on B200 GPUs

  • Multi-node inference

  • Fractional H100 GPUs for efficient model serving

RESPONSIBILITIES

  • Build and maintain scalable infrastructure to support the deployment and operation of machine learning models.

  • Establish standards and best practices for reliability and performance across the infrastructure.

  • Automate processes when relevant, particularly for managing CI/CD pipelines.

  • Own products and projects end-to-end, functioning as both an engineer and a project manager, with a focus on user empathy, project specification, and end-to-end execution.

  • Collaborate with cross-functional teams to understand project requirements and translate them into technical solutions.

  • Mentor junior team members and contribute to knowledge sharing within the organization.

  • Navigate ambiguity and exercise good judgment on tradeoffs and tools needed to solve problems, avoiding unnecessary complexity.

  • Demonstrate pride, ownership, and accountability for your work, expecting the same from your teammates.

REQUIREMENTS

  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field.

  • Extensive experience with Kubernetes.

  • Experience in building and maintaining scalable infrastructure.

  • Experience with infrastructure-as-code tools (e.g., Terraform, CloudFormation, Pulumi) and CI/CD tooling (e.g., GitHub Actions, GitLab CI, Circle CI, Jenkins).

  • Relevant OSS observability experience (Prometheus, ELK stack, Grafana stack, OpenTelemetry) is a plus.

  • Ability to own projects end-to-end, from project specification to execution.

  • No prior machine learning experience required, but should be open to learning about it.

BENEFITS

  • Competitive compensation, including meaningful equity

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

  • Paid parental leave

  • Fertility and family-building stipend through Carrot

  • Company-facilitated 401(k)

  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Similar Jobs

19 Days Ago
Remote
Canada
Senior level
Senior level
Database • Analytics
Design, build, and operate high-scale distributed observability systems for telemetry ingestion, processing, storage, autoscaling, and service provisioning. Own platform reliability, performance, capacity, and cost efficiency; participate in on-call incident response; automate repetitive operations; resolve root causes; and contribute to architecture, roadmap, and engineering quality across teams.
Top Skills: Argo CdAWSAzureClickhouseGCPGoGrafanaHelmKubernetesOpentelemetryPrometheusTerraformTypescript
14 Days Ago
In-Office or Remote
Senior level
Senior level
Healthtech • Information Technology
Design, deploy, and operate production-grade AKS clusters, owning cluster architecture, networking, security hardening, observability, CI/CD, and incident response. Integrate AKS with Azure services, manage deployments with Docker/Helm and GitOps, run Terraform IaC, define SLIs/SLOs, perform vulnerability management and DR planning, and provide on-call escalation and runbook documentation.
Top Skills: Application InsightsArgocdAzure CniAzure Container Registry (Acr)Azure DevopsAzure Key VaultAzure Kubernetes Service (Aks)Azure MonitorAzure Monitor For ContainersAzure PolicyAzure SqlAzure StorageBashCsi DriverDockerEvent HubsFluxGitopsGrafanaHelmIngress ControllersKafkaKubernetesLog Analytics WorkspaceLokiManaged IdentitiesNginxPrivate ClustersPrometheusPythonRbacTempoTerraformTraefik
2 Days Ago
Remote
CA
Senior level
Senior level
Artificial Intelligence • Information Technology • Software • Cybersecurity
Own and operate the data plane: develop Envoy/WASM filters for deep packet inspection, deploy and manage EKS workloads with Helm/GitOps and CI/CD, integrate Redis sidecars, provide observability (Prometheus/Grafana/OpenTelemetry), ensure reliable audit log streaming, and author ADRs as the data plane evolves.
Top Skills: Aws (EksCi/CdEksEnvoyGitopsGoGrafanaGrpcHelmIamKinesis)KmsKubernetesMtlsOpentelemetryPrometheusProtocol BuffersRedisRustSpiffeSpireTcp/IpWebassembly (Wasm)

What you need to know about the Montreal Tech Scene

With roots dating back to 1642, Montreal is often recognized for its French-inspired architecture and cobblestone streets lined with traditional shops and cafés. But what truly sets the city apart is how it blends its rich tradition with a modern edge, reflected in its evolving skyline and fast-growing tech industry. According to economic promotion agency Montréal International, the city ranks among the top in North America to invest in artificial intelligence, making it le spot idéal for job seekers who want the best of both worlds.

Key Facts About Montreal Tech

  • Number of Tech Workers: 255,000+ (2024, Tourisme Montréal)
  • Major Tech Employers: SAP, Google, Microsoft, Cisco
  • Key Industries: Artificial intelligence, machine learning, cybersecurity, cloud computing, web development
  • Funding Landscape: $1.47 billion in venture capital funding in 2024 (BetaKit)
  • Notable Investors: CIBC Innovation Banking, BDC Capital, Investissement Québec, Fonds de solidarité FTQ
  • Research Centers and Universities: McGill University, Université de Montréal, Concordia University, Mila Quebec, ÉTS Montréal

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account