Build and operate scalable, resilient, distributed services across AWS and Azure. The Senior Site Reliability Developer will automate infrastructure, develop platforms and frameworks, establish monitoring and incident response practices, define runbooks, reduce operational toil, lead delivery projects, mentor engineers, and participate in technical interviews and on-call rotations. The role requires expertise in cloud infrastructure, infrastructure as code, CI/CD, Kubernetes, observability, telemetry, and SRE practices including SLOs, SLIs, and error budgets.
This is a flexible position and has the option of working in our Toronto office full time, hybrid throughout the week or working remotely within Canada. Preference will be given to candidates located in the GTA who are able to attend the Toronto office approximately two days per month.
Vena is looking for a Senior SRE to join our SaaS Technology and Operations (STO) team. This role is a match for you if you love building highly scalable, resilient and automated services. We are an innovative team which aims to provide exceptional customer experience by leveraging best-in-class automation and orchestration practices for Vena's SaaS platform. As a Senior Site Reliability Developer, you will utilize your software and systems engineering background to build and run large-scale, distributed, fault-tolerant systems and services across AWS and Azure. We strive to hire people who are looking to make an impact and thrive in a flexible work environment driven by business objectives. Your role is to ensure that our systems - both internally and externally facing-have been designed with maximizing resiliency and uptime. Our team focuses on optimizing existing systems, building infrastructure and reducing toil through automation. Practices such as limiting time spent on manual operational work, post-mortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and technical standards.
Vena is looking for a Senior SRE to join our SaaS Technology and Operations (STO) team. This role is a match for you if you love building highly scalable, resilient and automated services. We are an innovative team which aims to provide exceptional customer experience by leveraging best-in-class automation and orchestration practices for Vena's SaaS platform. As a Senior Site Reliability Developer, you will utilize your software and systems engineering background to build and run large-scale, distributed, fault-tolerant systems and services across AWS and Azure. We strive to hire people who are looking to make an impact and thrive in a flexible work environment driven by business objectives. Your role is to ensure that our systems - both internally and externally facing-have been designed with maximizing resiliency and uptime. Our team focuses on optimizing existing systems, building infrastructure and reducing toil through automation. Practices such as limiting time spent on manual operational work, post-mortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and technical standards.
How You'll Make an Impact
- ·Helping Vena's technology organization build scalable systems, using best practices around automation (reliability) and developer self-service (velocity).
- Support services before they go live through activities such as system design consulting, developing software platforms and frameworks, planning and reviews. · Define and document runbooks and standard operating procedures.
- You will own and lead projects as part of the STO team’s delivery strategy.
- Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
- Provide mentorship and training to other Vena SREs as well as members of the Product and Technology organization on emerging technologies and new processes, drive education and knowledge transfer of design patterns and technical practices.
- Drive high standards around incident response practices and policies with a focus on automated response and remediation.
- Participate in influencing and shaping the overall STO team culture.
- You are a subject-matter expert in one or more technologies leveraged by the Vena platform.
- Participate in technical interviews for technical positions within STO and occasionally extend into other positions within Vena’s product and technology organization.
- Participate in on-call rotation.
What we use:
Please note this reflects only a portion of our current technical stack, and we are constantly evolving and revisiting our stack as we grow:
Please note this reflects only a portion of our current technical stack, and we are constantly evolving and revisiting our stack as we grow:
- A modern multi-cloud infrastructure across AWS and Azure, managed through infrastructure-as-code (Terraform) and configuration-as-code (Ansible)
- CI/CD primarily through Azure DevOps, with Jenkins still supporting some legacy pipelines
- Containerized workloads on Amazon EKS, AWS ECS, and Azure Container Apps
- RDS MySQL, Redshift, Redshift Spectrum, MongoDB, and Elasticsearch
- Kinesis, SQS, and RabbitMQ (via CloudAMQP)
- Workflow orchestration with Temporal
- Identity and access management with Auth0
- DevOps tools written in Python
- Back-end applications written using Java, Dropwizard, Spring Boot, and Hibernate
- Front-end applications written using TypeScript, JavaScript, React (Context Api and Hooks), and Redux
- Observability and monitoring with Observe Inc. (via OpenTelemetry), CloudWatch, and Azure Monitor
We'd Love to See
- 6+ years of experience in a Site Reliability Engineer role.
- You ideally possess an Associate or Professional level certification from AWS or Microsoft Azure.
- You are adept at core SRE concepts such as SLO, SLIs and error budgets and have direct experience in implementing them.
- In-depth knowledge of cloud computing platforms (AWS and Azure) and solid experience of setup and management of cloud infrastructure using IaC and orchestration tools.
- You can write code - in any language. You have implemented your work in a production environment and can back it up with examples.
- In-depth experience with tools and platforms such as: AWS, Azure, Ansible, Artifact storage (such as Artifactory, ECR), Build/Release Pipelines (such as Azure DevOps, Jenkins, Gitlab, GH Actions, or equivalents), Docker, Github, Kubernetes, Terraform etc.
- Direct experience with large-scale distributed systems in the cloud using observability and telemetry for oversight of code deployments.
- Experience with the operational aspects of software systems using telemetry, centralized logging, and alerting with tools such as: CloudWatch, Observe Inc, Prometheus, etc.
The base salary range for this position is $123,250 - 166,750 CAD.
*Our salaries are tailored to roles, levels and locations. Your individual pay within this range is influenced by factors like work location, skills, experience and education. As you progress in your role, your compensation may adapt, offering flexibility for growth beyond initial levels. For specifics, your recruiter will provide details and address any questions during the hiring process.
About
At Vena, we’re reimagining how businesses plan and grow: powered by data, collaboration, and innovation 🌱 Headquartered in Toronto with a global reach, we help finance teams work smarter using the tools they already love (yes, we’re talking about Excel!). Trusted by over 1,800 organizations and 150,000+ users every day, Vena Solutions brings clarity to complexity. We're a people led company that leads with our CORE values (Customer Trust, One Team, Respect & Authenticity and Execution Excellence) and are driven by our mission We’re growing fast, thinking big, and having fun along the way. The future of finance is being built here, and we’d love for you to be part of it 🚀Why Choose Us:💰 CompensationWe offer competitive and comprehensive total rewards packages that we review yearly to stay ahead of the market! Transparency is key, we keep you in the loop on how we design our comp programs. Build your future with our Employee Stock Option Program, Retirement Savings, Support & 401k Matching Programs.🎓 Personal Growth & LearningLevel up your skills with Vena! We support your journey with education subsidies, professional development programs, and learning opportunities to help you grow your career and reach your goals. Your future is bright and we’re here to invest in it!🧘♀️ Health & WellnessYour well-being = our priority! Great health & dental plans, wellness sessions (virtual & in-person), Employee Assistance Program (EAP), and a free Headspace subscription to support your mental health.🌍 Global ReachVena is everywhere you are! With offices in Toronto, London, and Indore, as well as team members across the United States, EMEA and beyond, we're a truly global company. Collaborate with colleagues around the world and bring diverse ideas to life: no passport required!🌴 Time OffRecharge with generous leave options - perfect for vacation, wellness, personal time, parental support, volunteering, and more.🏡 FlexibilityWe embrace a flexible culture that supports different ways of working, depending on the role and location. While some of our teams enjoy hybrid arrangements, others thrive in-office or remotely. With modern workspaces in Toronto, Indore, and London, we offer inspiring environments when you need them - and the freedom to work in ways that work best for you and your team.✨ AI-Powered InnovationAt Vena, we don’t just keep up, we're leading the way! 🚀We empower team members at every level to actively explore AI best practices, identify meaningful opportunities to apply AI in their work, and champion innovative solutions that drive impact.By fostering a culture of curiosity, innovation, and responsible use, we’re building a future where AI enhances how we work, think, and lead. From automating workflows to enhancing insights, AI is woven into everything we do.While every hiring decision is made by people, we may use AI to support application review.👉 Check us out:▶️YouTube💼 LinkedIn🔗GlassDoor* Vacancy Disclosure: Unless otherwise stated, all job postings on our careers site are for existing vacancies. Where explicitly stated, certain postings may be for anticipated vacancies that we expect to open in the near future.
Similar Jobs
Software
Operate and maintain highly available, secure, containerized SaaS applications across AWS and Azure. Responsibilities include rotating 24x7 on-call coverage, observability, incident and security response, disaster recovery, Terraform infrastructure-as-code, CI/CD automation, cloud integration, performance optimization, and developing AI agents to automate SRE and DevSecOps workflows.
Top Skills:
Amazon Web ServicesCheckovClaude CodeDatadogDockerDynatraceGithub ActionsGitlab CiGoJavaScriptKubernetesAzureNew RelicPrisma CloudPythonTerraformWiz
Cloud • Security • Software • Generative AI
Own and improve Elastic’s observability infrastructure across hosted cloud deployments. Responsibilities include Terraform-based infrastructure delivery, Python and Go development, production operations, incident response, on-call participation, RCA and postmortem writing, code and design reviews, mentoring, and improving operational documentation and processes. The role also involves operating Linux and containerized workloads, delivering complex projects independently, and maintaining secure, reliable platform infrastructure.
Top Skills:
AnsibleArgocdBeatsElastic Cloud Enterprise (Ece)Elastic Cloud Hosted (Ech)Elastic Cloud On Kubernetes (Eck)ElasticsearchGoHelmKibanaKubernetesKyvernoLinuxLogstashPuppetPythonTeleportTerraformVault
Software • Automation
Owns reliability, scalability, observability, and incident response for a mission-critical SaaS platform. Responsibilities include 24x7 on-call support, root cause analysis, automation, AWS infrastructure design, EKS/Kubernetes and Docker operations, Terraform and Helm deployments, CI/CD maintenance, cloud networking, monitoring with Datadog, Grafana, and Prometheus, database operations, and migration toward Kubernetes. The role also develops self-healing systems, documentation, runbooks, and cross-functional customer-focused reliability practices.
Top Skills:
AlbAmazon RdsAWSBashCi/CdDatadogDevsecopsDockerDocker SwarmEksGrafanaHelmIamInfrastructure As CodeKubernetesLinuxNlbPostgresPrometheusPythonRoute 53TerraformTerraformTransit GatewayVpcVpn
What you need to know about the Montreal Tech Scene
With roots dating back to 1642, Montreal is often recognized for its French-inspired architecture and cobblestone streets lined with traditional shops and cafés. But what truly sets the city apart is how it blends its rich tradition with a modern edge, reflected in its evolving skyline and fast-growing tech industry. According to economic promotion agency Montréal International, the city ranks among the top in North America to invest in artificial intelligence, making it le spot idéal for job seekers who want the best of both worlds.
Key Facts About Montreal Tech
- Number of Tech Workers: 255,000+ (2024, Tourisme Montréal)
- Major Tech Employers: SAP, Google, Microsoft, Cisco
- Key Industries: Artificial intelligence, machine learning, cybersecurity, cloud computing, web development
- Funding Landscape: $1.47 billion in venture capital funding in 2024 (BetaKit)
- Notable Investors: CIBC Innovation Banking, BDC Capital, Investissement Québec, Fonds de solidarité FTQ
- Research Centers and Universities: McGill University, Université de Montréal, Concordia University, Mila Quebec, ÉTS Montréal


