Baseten Logo

Baseten

Engineering Manager - Forward Deployed Engineering (LLM)

Reposted 22 Days Ago
Be an Early Applicant
Hybrid
Montréal, QC, CAN
Mid level
Hybrid
Montréal, QC, CAN
Mid level
Lead and mentor a team of Forward Deployed Engineers in building and optimizing LLM inference workloads, delivering high-performance AI applications, and collaborating with cross-functional teams.
The summary above was generated by AI

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products.

THE ROLE

As an Engineering Manager (Player & Coach), you will lead and mentor a team of Forward Deployed Engineers focused on building, scaling, and optimizing LLM inference workloads for Baseten customers. Applying both hands-on technical ownership and managerial leadership, you will guide your team through the processes of designing, deploying, and managing high performance, low latency AI applications on Baseten’s platform. FDE at baseten is not a sales function – we are a mix of engineering, product, and customer architects who contribute to the core Baseten codebase, drive large portions of our feature roadmap, and execute on complicated customer engagements.
You will also partner with product, infrastructure, and other customer engineering teams to ensure that large language models (LLMs) and other generative AI systems deliver best-in-class performance, reliability, and cost efficiency in production environments.
EXAMPLE INITIATIVES

Take a look at these blog posts written by members of our Forward Deployed Engineering team:

  • Forward Deployed Engineering on the frontier of AI

  • The fastest, most accurate Whisper transcription

  • Deploy production-ready model servers from Docker images

  • Deploy custom ComfyUI workflows as APIs

RESPONSIBILITIES

Leadership & Team Management

  • Lead, mentor, and grow a team of Forward Deployed Engineers, providing guidance on technical direction, project execution, and professional development.

  • Set clear goals and ensure timely, high-quality delivery across multiple customer-facing projects involving LLM deployment and inference optimization.

  • Collaborate with leadership to align team priorities with company and customer goals, balancing short-term delivery, widely varying customer priorities, and long-term technical initiatives.

  • Player-coach – While much of this role will be leading the team, you will also be expected to be a key driver on strategic product initiatives and customer engagements. The best managers derive credibility from being able to be hands-on when needed.

Technical Ownership

  • Develop and maintain software systems and product features using one or more general-purpose programming languages in a production-level environment, with a preference for Python due to its relevance in ML projects.

  • Drive customer impact by designing, implementing, and deploying Baseten solutions end-to-end (problem framing → evaluation → production deployment → monitoring). This involves working with customers’ engineering teams at every stage of the customer journey including: sales, implementation, and expansion.

  • Deliver with velocity: turn vague objectives into clear specs and well-defined PoCs so we can rapidly ship well-tested services and outcomes for our customers

  • Optimize and enhance AI/ML projects, contributing to the continuous improvement of our technical stack. This includes developing features and PRDs with other engineering and product orgs.

  • Own products and customer projects end-to-end, functioning as both an engineer, project manager, and product manager, with a focus on user empathy, project specification, and end-to-end execution.

REQUIREMENTS

  • Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or related field.

  • 4+ years of professional software engineering experience, including 1+ year in a leadership or mentorship capacity.

  • Strong programming skills in Python, with production experience in building or optimizing ML inference systems.

  • Proven experience with LLMs, inference optimization, or serving frameworks (e.g., vLLM, TensorRT, Triton, Hugging Face, Ray Serve).

  • Familiarity with observability, profiling, and cost/performance tradeoffs in production ML systems.

  • Excellent communication and collaboration skills—able to lead cross-functional efforts and drive outcomes in ambiguous, fast-paced environments.

BONUS POINTS

  • Experience leading customer-facing engineering teams or working directly with enterprise partners.

  • Deep understanding of GPU infrastructure, distributed inference, or model compression techniques.

BENEFITS

  • Competitive compensation, including meaningful equity.

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

  • Paid parental leave

  • Fertility and family-building stipend through Carrot

  • Company-facilitated 401(k)

  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Similar Jobs

Junior
eCommerce • Fashion • Retail • Sales • Wearables • Design
Deliver warm, professional customer service at cash wrap and on the sales floor. Operate POS, handle cash/media, receive and process shipments, maintain stockroom and inventory, replenish merchandise, execute visual merchandising updates, support loss prevention, and work flexible retail schedules with occasional heavy lifting.
Top Skills: Cash Register SystemsInternetIpadMobile PosPosWalkie Talkie
3 Hours Ago
In-Office or Remote
CA
Expert/Leader
Expert/Leader
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Senior Total Rewards partner advising on org design, pay-for-performance, equity, and executive compensation. Lead annual merit, equity, and promotion budgeting and execution, build dashboards and automation, prepare Compensation Committee/Board materials, and collaborate with HR, Finance, Legal, and Talent Acquisition to drive compensation strategy across the company.
Top Skills: Compensation Planning SoftwareWorkday
3 Hours Ago
In-Office or Remote
CA
Senior level
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Lead design and operation of Block's real-time, event-driven data systems. Build and run high-throughput streaming pipelines, turn Kafka event streams into low-latency data products, improve reliability and observability, mentor engineers, and define platform architecture and standards for projections, third-party data proxy, and graph data capabilities.
Top Skills: AWSBlock-ForgeCloudeventsDatabricksDatahubDbtFlinkIcebergKafkaKotlinPrefectSnowflakeSparkSpark Structured Streaming

What you need to know about the Montreal Tech Scene

With roots dating back to 1642, Montreal is often recognized for its French-inspired architecture and cobblestone streets lined with traditional shops and cafés. But what truly sets the city apart is how it blends its rich tradition with a modern edge, reflected in its evolving skyline and fast-growing tech industry. According to economic promotion agency Montréal International, the city ranks among the top in North America to invest in artificial intelligence, making it le spot idéal for job seekers who want the best of both worlds.

Key Facts About Montreal Tech

  • Number of Tech Workers: 255,000+ (2024, Tourisme Montréal)
  • Major Tech Employers: SAP, Google, Microsoft, Cisco
  • Key Industries: Artificial intelligence, machine learning, cybersecurity, cloud computing, web development
  • Funding Landscape: $1.47 billion in venture capital funding in 2024 (BetaKit)
  • Notable Investors: CIBC Innovation Banking, BDC Capital, Investissement Québec, Fonds de solidarité FTQ
  • Research Centers and Universities: McGill University, Université de Montréal, Concordia University, Mila Quebec, ÉTS Montréal

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account