Fabrion Logo

Fabrion

DevOps Engineer (Founding Team)

Reposted 2 Months Ago
In-Office or Remote
Hiring Remotely in CA
Senior level
In-Office or Remote
Hiring Remotely in CA
Senior level
Design, build, and operate secure, tenant-isolated cloud infrastructure and CI/CD for an AI-native multi-tenant platform. Implement observability, policy-as-code, SLAs, incident response, and automate provisioning (Terraform/Helm). Support deployment, monitoring, and secure operations for ML/agent systems at scale.
The summary above was generated by AI

DevOps Engineer (Founding Team)

Location: San Francisco Bay Area

Type: Full-Time

Compensation: Competitive salary + meaningful equity (founding tier)

Backed by 8VC, we're building a world-class team to tackle one of the industry’s most critical infrastructure problems.

About the Role

We're building an AI-native, multi-tenant enterprise platform for complex domains in industrial verticals. In this architecture, DevOps isn't just about shipping features — it's about operationalizing intelligent agents, ensuring traceability across AI systems, and supporting mission-critical ML infrastructure at scale.

We're looking for a DevOps engineer who can own infrastructure from Day 1 — automating everything from CI/CD and observability to cloud governance and security. You’ll work with a highly technical team building real-time AI pipelines and multi-agent systems. If you want to be the person who makes the platform run — fast, secure, reliable, and explainable — this is your role.

Responsibilities
  • Build and maintain scalable cloud infrastructure across AWS/GCP/Azure with a focus on secure, tenant-isolated deployments

  • Own and evolve CI/CD systems (e.g. GitHub Actions, ArgoCD) with progressive rollout, testing, and rollback flows

  • Establish observability tooling across services, agents, and pipelines (OpenTelemetry, Prometheus, Grafana, Sentry)

  • Implement policy-as-code (OPA, Rego) for deployment safety, RBAC, audit logging, and approval workflows

  • Define and enforce SLAs, uptime targets (99.99%+), incident response, and remediation workflows

  • Secure infrastructure: IAM, VPC, encryption, key management, image scanning, secrets rotation

  • Automate deployments, infrastructure provisioning (Terraform, Helm), and environment replication

What We’re Looking For

Core Experience:

  • 4–10+ years in DevOps, platform engineering, or SRE in production-grade systems

  • Strong experience with Docker, Kubernetes (EKS/GKE), Terraform or Pulumi

  • Hands-on experience deploying and monitoring distributed cloud-native systems

  • Familiar with GitOps practices, CI/CD design, progressive delivery, and secure SDLC

  • Clear understanding of how to implement monitoring, alerting, and failure simulation in dynamic environments

Engineering Mindset:

  • Obsessed with reliability, latency, uptime, and repeatability

  • Security-aware and compliance-conscious

  • Proactive — you don’t wait for alerts to fix things

  • Comfortable collaborating with backend, AI, and data teams

Bonus: Agent-Native / ML Ops Capabilities

  • We’re building an agentic, AI-native platform from the ground up. Experience here isn’t required, but would be a strong differentiator:

  • Experience running LLM orchestration frameworks (e.g. LangChain, LangGraph, Dust, ReAct agents)

  • Building retrieval-augmented generation (RAG) pipelines — and deploying them safely and repeatably

  • Familiarity with vector DBs (Weaviate, Qdrant, Pinecone) and embedding pipelines

  • Monitoring and governing long-running or multi-agent chains

  • Auditability and replay systems for agent decision-making

  • Serving fine-tuned or open-source LLMs with model versioning and GPU scaling (e.g. vLLM, TGI)

  • Interest in auto-remediation using agents (e.g. observability + alert → insight → response via LLM)

Why This Role Matters

DevOps is the nervous system of the platform — every agent, every data fabric component, every pipeline flows through what you build. This is a rare opportunity to design that system early, the right way, and future-proof it for scale, compliance, and trust.

If you're excited by intelligent systems, distributed data, and deeply technical infrastructure problems — and you want your work to have immediate real-world impact — we’d love to hear from you.

Similar Jobs

2 Months Ago
In-Office or Remote
Mid level
Mid level
Blockchain • Financial Services • Cryptocurrency • Web3
The DevOps Engineer will enhance CI/CD workflows, manage AWS infrastructure, automate operations, and improve system reliability and monitoring.
Top Skills: AWSCeleryDatadogDockerEcsGithub ActionsGoGrafanaKubernetesPostgresPrometheusPulumiPythonRedisTerraform
Yesterday
Easy Apply
Remote or Hybrid
Canada
Easy Apply
Junior
Junior
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Serve as the primary post-implementation contact for top customers, craft joint success plans, run executive business reviews and workshops, mentor teammates, support French-speaking customers, and advise on customizing Samsara’s IoT platform to drive safety, efficiency, and sustainability.
Top Skills: IotSamsara PlatformVehicle TelematicsVideo-Based Safety
Yesterday
Easy Apply
Remote or Hybrid
Canada
Easy Apply
Expert/Leader
Expert/Leader
Marketing Tech • Social Media • Software • Analytics • Business Intelligence
Own the strategy, roadmap, delivery, pricing, and growth of Sprout Social’s Listening product. Drive adoption and competitive differentiation through customer discovery, market analysis, product investments, and cross-functional collaboration with Engineering, Design, Marketing, Sales, Customer Experience, and GTM teams. Establish success metrics, optimize usage and pricing opportunities, communicate product direction, and lead roadmap execution across distributed teams.
Top Skills: AnalyticsProduct AnalyticsSaaSSocial Media Intelligence

What you need to know about the Montreal Tech Scene

With roots dating back to 1642, Montreal is often recognized for its French-inspired architecture and cobblestone streets lined with traditional shops and cafés. But what truly sets the city apart is how it blends its rich tradition with a modern edge, reflected in its evolving skyline and fast-growing tech industry. According to economic promotion agency Montréal International, the city ranks among the top in North America to invest in artificial intelligence, making it le spot idéal for job seekers who want the best of both worlds.

Key Facts About Montreal Tech

  • Number of Tech Workers: 255,000+ (2024, Tourisme Montréal)
  • Major Tech Employers: SAP, Google, Microsoft, Cisco
  • Key Industries: Artificial intelligence, machine learning, cybersecurity, cloud computing, web development
  • Funding Landscape: $1.47 billion in venture capital funding in 2024 (BetaKit)
  • Notable Investors: CIBC Innovation Banking, BDC Capital, Investissement Québec, Fonds de solidarité FTQ
  • Research Centers and Universities: McGill University, Université de Montréal, Concordia University, Mila Quebec, ÉTS Montréal

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account