LawZero Logo

LawZero

Senior Data Platform Engineer

Posted 10 Days Ago
Be an Early Applicant
In-Office
Montréal, QC, CAN
Senior level
In-Office
Montréal, QC, CAN
Senior level
Architect, implement, scale, and maintain petabyte-scale data processing platforms for frontier AI models. Build reliable, observable, cost-efficient pipelines with heterogeneous and GPU-accelerated workloads, including failure recovery and restartability. Ensure datasets and intermediate outputs are versioned, reproducible, traceable, and documented with datasheets. Partner with AI researchers to integrate datasets into training pipelines and guide sustainable technology decisions.
The summary above was generated by AI

We are seeking a visionary and highly technical Senior Data Platform Engineer to architect, implement, scale, and maintain the data engine powering our next-generation frontier models.

In this high-impact role, you will bridge the gap between cutting-edge AI research and high-performance engineering, treating the data platform as an internal product with our researchers as your primary customers. You will be responsible for building automated, petabyte-scale data processing pipelines and for guaranteeing that everything they produce is efficient, reliable, and fully traceable. Our technical environment is not fixed and will evolve as our projects scale. We expect someone capable of evolving it, not only following industry trends, challenging it, and making sustainable decisions in close collaboration with our Research and Product teams.

The title of Engineer is used for reference purposes and may or may not be the official title of the applicant based on jurisdiction.

Key Responsibilities
  • Scale and automate the data processing stack to handle petabytes of data and ensure its smooth operation.
  • Design the execution layer for pipeline stages of differing computational shape, matching each stage to an appropriate engine and keeping the pipeline saturated end to end.
  • Ensure efficient use of compute resources, including GPU access for compute-intensive data processing tasks.
  • Make pipelines reliable at scale, with graceful failure recovery, restartability, and observability that attributes bottlenecks and cost to the responsible stage.
  • Ensure all datasets, including the intermediate outputs of each transformation stage, are versioned, reproducible, and fully traceable to meet specific and dynamic experiment needs, and are accompanied by datasheets, in accordance with internal Data Governance policies.
  • Partner with the Research team to ensure the datasets you produce integrate seamlessly with the training pipelines.
Skills and Qualifications
  • A bachelor's degree in a relevant field (e.g., computer science, computer engineering, software engineering) is required.
  • 5+ years of experience designing, implementing, and managing large-scale distributed data processing systems, or working within large-scale distributed ML data frameworks, with recent experience using e.g. Ray, Apache Spark, workflow orchestrators, Apache Arrow, and/or Parquet.
  • Demonstrated ownership of a data processing system under real throughput, reliability, and cost pressure. The specific frameworks matter less to us than evidence that you have had to reason about where a large pipeline breaks and why.
  • Experience profiling and optimizing throughput and cost across heterogeneous workloads, including GPU-accelerated stages.
  • Experience with dataset versioning, lineage, and reproducibility tooling.
  • Ability to collaborate effectively with cross-functional teams, document best practices, and stay updated with the latest advancements in large-scale data processing and software development.
  • Experience with workload managers (e.g., Ray, Kubernetes, Slurm).
  • Familiarity with containerization tools (e.g., Docker, Enroot).
  • Familiarity with data infrastructures and platforms (e.g., vector databases).
What we offer
  • The chance to contribute meaningfully to a globally critical initiative.
  • Comprehensive health benefits (including mental health and wellness management account).
  • 20 days of vacation per year upon start.
  • Employer contribution of 4% to your retirement savings, with no required employee match.
  • Additional compensation totalling 8% of your salary to apply towards additional retirement savings or bonuses (independent of group and individual performance).
  • A team of passionate world-class experts in their field.
  • A collaborative and inclusive work environment in our vibrant office space in the heart of Little Italy, in the trendy Mile-Ex district, close to public transportation.

About LawZero

LawZero is a non-profit organization committed to advancing research and creating technical solutions that enable safe-by-design AI systems. Its scientific direction is based on new research and methods proposed by Professor Yoshua Bengio, the most cited AI researcher in the world. Based in Montreal, LawZero’s research aims to build non-agentic AI that learns primarily to understand the world rather than to act in it, giving truthful answers to questions based on transparent and externalized probabilistic reasoning. Such AI systems could be used to accelerate scientific discovery, to provide oversight for agentic AI systems, and to advance the understanding of AI risks and how to avoid them. LawZero believes that AI should be cultivated as a global public good—developed and used safely towards human flourishing. For more information, visit www.lawzero.org

You belong here

At LawZero, diversity is important to us. We value a work environment that is fair, open and respectful of differences. We welcome applications from highly qualified individuals interested in working towards our mission in a respectful, inclusive and collaborative setting.

Your personal information will be collected and processed by LawZero to evaluate your application for employment in compliance with our Privacy Policy. Under privacy laws in force in your country of residence, you may have several privacy rights, such as to request access to your personal information or to request that your personal information be rectified or erased. Details on how you can exercise your rights can be found in our Privacy Policy.

Similar Jobs

24 Days Ago
Easy Apply
Hybrid
Easy Apply
Senior level
Senior level
AdTech • Artificial Intelligence • Digital Media • Marketing Tech • Social Media • Software • Generative AI
Design and operate a governed data platform for contract-driven publishing and consumption. Build PostgreSQL-to-Debezium-to-Kafka CDC workflows, GitOps-style contract validation and reconciliation, developer tooling, Kubernetes runtimes, and DataHub catalog integrations. Define data product standards covering schemas, ownership, lifecycle, compatibility, lineage, access, and recovery. Collaborate with infrastructure, product, analytics, data engineering, and AI/ML teams to deliver reliable, reusable production data capabilities.
Top Skills: Amazon S3Apache IcebergAvroBigQueryClickhouseDatahubDebeziumGithub ActionsGitopsGoInfrastructure-As-CodeJson SchemaKafkaKubernetesPolicy-As-CodePostgresProtobufPythonRubyTypescript
18 Days Ago
In-Office
Senior level
Senior level
Fintech • Payments • Software • Financial Services
Own and scale the data platform supporting multi-tenant industrial optimization applications. Responsibilities include ingesting heterogeneous customer data, designing storage and query paths for millions of records, managing PostgreSQL performance and migrations, building infrastructure and deployment systems, improving optimization-engine performance, and shipping and operating features in production. The role works across Python, TypeScript, cloud infrastructure, APIs, data pipelines, and the broader product stack.
Top Skills: Apache ArrowCi/CdCloudflareCsvDbtDuckdbExcelHTTPJIRALinearParquetPostgresPythonS3Solid.JsSQLTypescript
Senior level
eCommerce • Events • News + Entertainment • Software
Design, build, and optimize real-time and batch data pipelines using Apache Beam and GCP services. Integrate multiple data sources, ensure data quality and observability, improve pipeline performance and cost efficiency, and apply CI/CD, infrastructure-as-code, monitoring, governance, security, and compliance practices. Collaborate with data scientists, analysts, and software engineers to deliver reliable data for analytics, AI, and business intelligence.
Top Skills: Apache BeamBigQueryCi/CdCloud ComposerCloud Deployment ManagerCloud StorageDatadogDataflowFlinkGoogle Cloud Platform (Gcp)OpenlineagePrometheusPub/SubPythonSpark Structured StreamingTerraform

What you need to know about the Montreal Tech Scene

With roots dating back to 1642, Montreal is often recognized for its French-inspired architecture and cobblestone streets lined with traditional shops and cafés. But what truly sets the city apart is how it blends its rich tradition with a modern edge, reflected in its evolving skyline and fast-growing tech industry. According to economic promotion agency Montréal International, the city ranks among the top in North America to invest in artificial intelligence, making it le spot idéal for job seekers who want the best of both worlds.

Key Facts About Montreal Tech

  • Number of Tech Workers: 255,000+ (2024, Tourisme Montréal)
  • Major Tech Employers: SAP, Google, Microsoft, Cisco
  • Key Industries: Artificial intelligence, machine learning, cybersecurity, cloud computing, web development
  • Funding Landscape: $1.47 billion in venture capital funding in 2024 (BetaKit)
  • Notable Investors: CIBC Innovation Banking, BDC Capital, Investissement Québec, Fonds de solidarité FTQ
  • Research Centers and Universities: McGill University, Université de Montréal, Concordia University, Mila Quebec, ÉTS Montréal

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account