Arena (arena.ai) Logo

Arena (arena.ai)

Machine Learning Engineer

Reposted 10 Days Ago
Remote or Hybrid
Hiring Remotely in CA
Senior level
Remote or Hybrid
Hiring Remotely in CA
Senior level
Senior Machine Learning Engineer role focusing on building and improving AI evaluation infrastructures, collaborating with a multidisciplinary team, and contributing to the development of a scalable evaluation platform.
The summary above was generated by AI
About Arena Intelligence

Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.


Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.


We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.

About the Role

Arena Intelligence is seeking a Senior Machine Learning Engineer to help scale and strengthen the core infrastructure that powers real-world AI evaluation. You’ll play a foundational role in shaping how we build, deploy, and improve our model benchmarking systems, working across data pipelines, inference APIs, and new evaluation methodologies. This is an opportunity to apply your technical expertise to a platform trusted by millions, and to help define how cutting-edge AI is assessed in the wild.

As one of the first ML engineers on the team, you’ll partner closely with researchers, engineers, and product leadership to turn new ideas into reliable systems. You’ll help us move fast while staying rigorous, improving reproducibility, scaling up to new modalities, and deepening our ability to understand and compare frontier models.

You’ll
  • Architect and build what will become our core modeling for data and evaluation products

  • Own the full stack data, model training, and eval pipelines

  • Help grow a culture of feedback and rapid product iteration as we build new features as a tight-nit team

  • Conduct research into state-of-the-art evaluation methods and contribute to the long-term vision for a centralized, scalable evaluation platform.

You’ll have
  • Strong programming skills with the ability to work across the stack in a typical recommendation system or LLM stack

  • Experience in deep learning, language models or reward model training

  • Experience in working with LLM for fine tuning, prompt engineering, function calling etc

  • Self-motivated with a willingness to take ownership of tasks

  • A passion for shipping quality products

  • 4+ years of industry experience or relevant projects

  • Solid understanding of statistics, and various tools and methodologies for evaluating uncertainty in a way that is specific to the given product being shipped

What we offer
  • We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.

  • Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.

  • The opportunity to work on cutting-edge AI with a small, mission-driven team

  • A culture that values transparency, trust, and community impact

Come help build the space where anyone can explore and help shape the future of AI.

Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.

Similar Jobs

3 Days Ago
In-Office or Remote
CA
Expert/Leader
Expert/Leader
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Design, build, and operate production ML decision systems to detect and prevent payment fraud, account takeover, scams, and other abuse. Integrate diverse signals into low-latency serving and batch scoring, own feature pipelines and model lifecycle, develop AI-assisted triage and feedback loops, and partner cross-functionally to balance fraud reduction with legitimate customer access.
Top Skills: Cloud InfrastructureData LakehouseData WarehouseEmbeddingsFeature StoreJavaKafkaKotlinKubernetesLightgbmModel ServingMonitoringObservabilityPythonPyTorchSQLTensorFlowWorkflow OrchestrationXgboost
3 Days Ago
In-Office or Remote
CA
Expert/Leader
Expert/Leader
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Design, build, and operate production ML systems that generate trusted signals for ranking, retrieval, recommendations, propensity/churn/LTV, and next-best-action decisioning. Define signal/data contracts, own feature and candidate generation through serving, experimentation, monitoring, and feedback loops, and evaluate long-term business impact, trust, fairness, and compliance. Partner across product, data, modeling, risk, and compliance and apply AI/agents to accelerate engineering and operations.
Top Skills: Agent-Assisted Operations ToolingBatch PipelinesCloud InfrastructureCoding AgentsData WarehousesEmbeddingsEvaluation HarnessesEvent StreamsExperimentation SystemsFeature StoresJavaKotlinKubernetesLakehousesLightgbmModel-Serving InfrastructureObservability ToolingPythonPyTorchRanking/Retrieval SystemsRecommendation FrameworksSemantic SearchSQLTensorFlowWorkflow OrchestrationXgboost
3 Days Ago
Remote or Hybrid
CA
Expert/Leader
Expert/Leader
Blockchain • Fintech • Mobile • Payments • Software • Financial Services
Design, build, and operate production ML decision systems to detect and prevent payment fraud, account takeover, identity abuse, merchant/marketplace risk, scams, and other adversarial activity. Own end-to-end production lifecycle: data contracts, low-latency inference, batch scoring, feature quality, deployment, monitoring, incident response, rollback, and feedback loops. Develop AI-assisted workflows and reusable decision capabilities while partnering with modelers, product, compliance, and operations.
Top Skills: Agent-Assisted ToolingCloud InfrastructureCoding AgentsData WarehouseEmbeddingsFeature StoreJavaKafkaKotlinKubernetesLakehouseLightgbmModel-Serving SystemsMonitoringObservabilityPythonPyTorchSQLTensorFlowWorkflow OrchestrationXgboost

What you need to know about the Montreal Tech Scene

With roots dating back to 1642, Montreal is often recognized for its French-inspired architecture and cobblestone streets lined with traditional shops and cafés. But what truly sets the city apart is how it blends its rich tradition with a modern edge, reflected in its evolving skyline and fast-growing tech industry. According to economic promotion agency Montréal International, the city ranks among the top in North America to invest in artificial intelligence, making it le spot idéal for job seekers who want the best of both worlds.

Key Facts About Montreal Tech

  • Number of Tech Workers: 255,000+ (2024, Tourisme Montréal)
  • Major Tech Employers: SAP, Google, Microsoft, Cisco
  • Key Industries: Artificial intelligence, machine learning, cybersecurity, cloud computing, web development
  • Funding Landscape: $1.47 billion in venture capital funding in 2024 (BetaKit)
  • Notable Investors: CIBC Innovation Banking, BDC Capital, Investissement Québec, Fonds de solidarité FTQ
  • Research Centers and Universities: McGill University, Université de Montréal, Concordia University, Mila Quebec, ÉTS Montréal

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account