Xsolla Logo

Xsolla

Site Reliability Engineer (Monetization)

Posted 6 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in Montréal, QC, CAN
Mid level
In-Office or Remote
Hiring Remotely in Montréal, QC, CAN
Mid level
Own application-level infrastructure for the Monetization domain: Helm/Terraform/Kubernetes, observability (SLOs/SLIs/Datadog/OpenTelemetry), CI/CD, capacity planning, incident response, runbooks, automation, and reliability roadmaps while embedded with product teams.
The summary above was generated by AI

ABOUT YOU

    We are looking for a Site Reliability Engineer (Monetization) who is pragmatic, product-minded, and equally comfortable writing code and running production systems to join our Infrastructure department's SRE team. The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to own the application-level infrastructure and reliability of a high-traffic commerce domain end to end - from deploy pipelines and Kubernetes manifests to SLOs, capacity planning, and production readiness.

    Strong Kubernetes, observability, and software engineering skills are essential, along with experience in operating production services in a cloud environment (GCP/GKE or comparable) and partnering closely with product development teams. The ability to hold a dual perspective - understanding both how developers ship features and what infrastructure needs to stay reliable - and to bring the reliability lens into design decisions early will be key to your success in this role.

    This is a hybrid embedded role: you remain part of the SRE organization (practices, standards, duty rotation) while being functionally embedded into the Monetization product domain. You'll build long-term working relationships with the domain's engineering teams, own a meaningful share of their application infrastructure execution, and co-author the reliability practices used company-wide.

    If you're passionate about making complex distributed systems boringly reliable and love building the commerce and monetization backbone that lets game developers around the world get paid, we would love to hear from you!

ABOUT US

    Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators. Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win, developers have all the things needed to enjoy the game.

    For more information, visit xsolla.com.

Responsibilities

  • Own the application-level infrastructure of the Monetization domain: Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, and service-level networking and integrations
  • Own the domain's observability: design and implement SLOs/SLIs, monitors, alerts, and dashboards for critical services on Datadog and OpenTelemetry-based tooling
  • Help to set up and evolve CI/CD pipelines for domain services (GitLab CI, GitHub Actions), including deploy and rollback automation
  • Perform capacity planning and performance tuning ahead of expected load - product launches, sales events, and regional rollouts - including load testing and performance regression investigation
  • Run Production Readiness Reviews for new services and major changes; define and enforce what "production-ready" means for the domain
  • Support domain incident response: assist with deep investigation of complex incidents, contribute to post-mortems, drive follow-up reliability improvements, and maintain runbooks
  • Build domain-specific automation that reduces operational toil: runbook automation, deploy helpers, recurring operational scripts
  • Maintain and drive a forward-looking reliability roadmap for the domain together with product engineering leads
  • Participate in product team planning, refinements, and architecture reviews, bringing the reliability perspective before design decisions become expensive to change
  • Co-author company-wide SLO/SLI, capacity, and operational standards together with the broader SRE team; contribute improvements directly to shared SRE-operated subsystems
  • Participate in the SRE duty rotation, supporting developers across the company

Qualifications & Skills

    • 3+ years of proven SRE, DevOps, or platform engineering experience: on-call or incident response duty, SLO/monitoring ownership, deploy pipeline and infrastructure work for production services
    • Software development background: you have built and shipped backend services, not only operated them - comfortable reading application code during an investigation and writing production-quality automation in at least one language (e.g., Go, PHP)
    • Hands-on Kubernetes experience: Helm, manifests, deploy strategies, debugging application-level performance and networking issues (GKE or another managed Kubernetes)
    • Solid observability practice: building monitors, dashboards, and SLOs/SLIs on a modern platform (Datadog preferred; Prometheus/Grafana experience also relevant), familiarity with OpenTelemetry
    • Infrastructure as Code exposure (Terraform/Terragrunt) for collaboration with platform teams
    • GCP experience (IAM, networking, managed services)
    • Experience building and maintaining CI/CD pipelines (GitLab CI and/or GitHub Actions)
    • Programming/scripting proficiency sufficient to build automation and tooling (e.g., Python, Go, or Bash)
    • Practical experience with incident response, post-mortems, and driving reliability improvements from incidents
    • Strong collaboration and communication skills — this role works embedded with product development teams daily
    • Experience in payments, fintech, e-commerce, or gaming — high-traffic transactional systems
    • Nice to Have:
    • Kubernetes certifications
    • Google Cloud Platform certifications
    • HashiCorp certifications

Xsolla Montréal, Québec, CAN Office

Montréal, Canada

Similar Jobs

2 Hours Ago
In-Office or Remote
CA
Expert/Leader
Expert/Leader
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Senior Total Rewards partner advising on org design, pay-for-performance, equity, and executive compensation. Lead annual merit, equity, and promotion budgeting and execution, build dashboards and automation, prepare Compensation Committee/Board materials, and collaborate with HR, Finance, Legal, and Talent Acquisition to drive compensation strategy across the company.
Top Skills: Compensation Planning SoftwareWorkday
2 Hours Ago
In-Office or Remote
CA
Senior level
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Lead design and operation of Block's real-time, event-driven data systems. Build and run high-throughput streaming pipelines, turn Kafka event streams into low-latency data products, improve reliability and observability, mentor engineers, and define platform architecture and standards for projections, third-party data proxy, and graph data capabilities.
Top Skills: AWSBlock-ForgeCloudeventsDatabricksDatahubDbtFlinkIcebergKafkaKotlinPrefectSnowflakeSparkSpark Structured Streaming
2 Hours Ago
In-Office or Remote
CA
Senior level
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Lead pricing modeling and experimentation to optimize margin, conversion, and retention. Build elasticity and willingness-to-pay models, run causal and bandit experiments, decompose merchant economics, and develop agentic deal tooling. Own end-to-end data pipelines, ETL, monitoring, and visualization. Evaluate AI systems for accuracy, bias, and drift, and partner with cross-functional teams to translate results into pricing strategy and automation.
Top Skills: Agentic AiLlmsLookerNumpyPandasPythonRRagSnowflakeSQLTableau

What you need to know about the Montreal Tech Scene

With roots dating back to 1642, Montreal is often recognized for its French-inspired architecture and cobblestone streets lined with traditional shops and cafés. But what truly sets the city apart is how it blends its rich tradition with a modern edge, reflected in its evolving skyline and fast-growing tech industry. According to economic promotion agency Montréal International, the city ranks among the top in North America to invest in artificial intelligence, making it le spot idéal for job seekers who want the best of both worlds.

Key Facts About Montreal Tech

  • Number of Tech Workers: 255,000+ (2024, Tourisme Montréal)
  • Major Tech Employers: SAP, Google, Microsoft, Cisco
  • Key Industries: Artificial intelligence, machine learning, cybersecurity, cloud computing, web development
  • Funding Landscape: $1.47 billion in venture capital funding in 2024 (BetaKit)
  • Notable Investors: CIBC Innovation Banking, BDC Capital, Investissement Québec, Fonds de solidarité FTQ
  • Research Centers and Universities: McGill University, Université de Montréal, Concordia University, Mila Quebec, ÉTS Montréal

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account