Velixo Logo

Velixo

Applied AI Engineer

Posted 9 Days Ago
Be an Early Applicant
In-Office
Montréal, QC, CAN
Entry level
In-Office
Montréal, QC, CAN
Entry level
Owns evaluation, observability, and cost optimization for Velixo Intelligence, an MCP-based AI gateway for ERP queries and governed writebacks. Responsibilities include building regression and evaluation suites, instrumenting token and cost metrics, analyzing model performance, optimizing prompts, caching, context, and routing, tracking model releases, and supplying unit-cost data for Finance and pricing decisions.
The summary above was generated by AI

About Velixo
Velixo builds Excel-native reporting, planning, and automation for finance teams running Acumatica, Sage Intacct, Business Central, and MYOB. Founded in 2017 and headquartered in Montreal, we serve finance and accounting teams who live in spreadsheets and need their ERP data to be live, trustworthy, and writable. 
Velixo is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
This is a fully remote position. While we welcome applications from across Canada, we are giving priority to candidates located in Quebec at this time.
The role 
You will own how well Velixo Intelligence works and what it costs us to run. That means building the evaluation infrastructure that tells us whether a model change made our product better or worse, and driving down the cost per action without degrading quality. 

This is a measurement-first engineering role. Today we ship AI features and form opinions about quality from anecdotes. You will replace that with evidence. 

About Velixo Intelligence 

Velixo Intelligence is our MCP-based gateway that lets finance teams query their ERP in natural language and perform governed writebacks from Claude, ChatGPT, and Copilot. Reads run through GIQL, our query language. Writes go through the ERP's own screen APIs so that every permission, validation, and audit rule still applies. 

It launched in public preview this year. Usage is growing, the surface area is expanding, and our understanding of the unit economics has not kept pace. That gap is the reason this role exists. 

What you will own 

Evaluation 

  • Build and maintain regression suites for GIQL query generation, writeback correctness, and tool selection 

  • Define what "good" means for each action type and make it measurable 

  • Run structured evaluations before every model upgrade or prompt change, and produce a clear ship or no-ship recommendation 

  • Track quality over time so we notice degradation before customers do 

Cost engineering 

  • Instrument every Gateway action for tokens, cache behavior, model, latency, and tenant 

  • Establish and maintain the cost-per-action baseline that the rest of the company plans against 

  • Reduce cost through prompt compression, prompt caching, context pruning, and routing work to smaller models where evaluation shows quality holds 

  • Identify the expensive tail: which customers, which query shapes, which failure and retry loops 

Model strategy 

  • Track new model releases across providers and evaluate them against our workloads rather than published benchmarks 

  • Maintain a current view of the price and performance frontier for what we do 

  • Recommend when to migrate, when to wait, and when to run models in parallel 

Partnership with Finance 

  • Supply the unit cost data behind our credit pricing and plan allowances 

  • Model the margin impact of proposed pricing changes alongside our VP Finance 

  • Flag when product decisions will move the cost curve before they ship 

You will not own pricing decisions. You will own the numbers those decisions depend on. 

First 90 days 

  • Days 1 to 30: Full instrumentation of the Gateway. A dashboard showing cost per action type, per tenant, and per model that the exec team checks weekly. 

  • Days 31 to 60: A working evaluation suite covering our highest-volume action types, with a documented baseline. 

  • Days 61 to 90: A prioritized cost reduction roadmap with sized estimates, and at least one shipped optimization with measured before-and-after quality. 

What we are looking for 

  • Substantial experience building on LLM APIs in production, not in demos or notebooks 

  • You have built an eval harness before and can explain why the naive version of it misleads you 

  • Hands-on with the current tooling landscape: tracing and observability platforms such as Langfuse, eval frameworks such as promptfoo. You have opinions about which of these earn their keep and which add ceremony without adding signal. 

  • Comfortable with token accounting, context window management, and caching strategies 

  • Strong analytical instincts and fluency with data. You reach for a query or a notebook before you reach for an opinion. 

  • Able to write clearly about tradeoffs for a non-technical audience, because your conclusions will land in board decks and pricing meetings 

  • Python or TypeScript proficiency 

Nice to have 

  • Experience with MCP or other tool-calling frameworks 

  • Background in ERP, accounting, or financial systems 

  • Prior exposure to usage-based or credit-based pricing models 

  • Experience at a small company where you set your own priorities 

Why this role is interesting 

Most companies bolt a chat box onto their product and hope. We are building governed, auditable write access to financial systems of record, where a wrong answer has real consequences and a hallucinated journal entry is unacceptable. The evaluation problem here is genuinely hard and genuinely matters. 

You will also have unusual visibility. The numbers you produce will directly shape what we charge and what we build. 

HQ

Velixo Montréal, Québec, CAN Office

2575 Place Chassé, Suite 200, Montréal, Quebec , Canada, H1Y 2C3

Similar Jobs

9 Days Ago
Remote or Hybrid
Canada
Senior level
Senior level
Big Data • Information Technology • Software • Database • Analytics
Develop experimental AI techniques for agentic marketing applications, especially image and video generation. Build proofs of concept, autonomous evaluation and improvement systems, realistic AI-generated video, and brand-aligned creative generation workflows. The role requires strong quantitative and probabilistic thinking, backend architecture skills, creativity with LLM applications, and product intuition.
Top Skills: Agentic AiBackend ArchitectureGenerative AiImage GenerationLlmsMachine LearningProbabilistic SystemsVideo Generation
17 Days Ago
Remote or Hybrid
Canada
Entry level
Entry level
Software
Own the generation side of Document Intelligence by designing prompts, schemas, validators, multimodal context assembly, and model-selection strategies. Build evaluation datasets and rubrics, monitor factuality and quality, identify regressions, and ship reliable LLM features into production. Work with video, PDF, audio, and image inputs to generate validated maintenance entities while balancing quality, cost, and latency.
Top Skills: LlmsLlmx Eval PipelineMultimodal AiOcrPrompt EngineeringRagRetrieval SystemsStructured OutputTool Use
23 Days Ago
In-Office or Remote
CA
Senior level
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Lead architecture and technical strategy for AI-driven product quality systems using LLMs and agents. Build scalable evaluation frameworks, detect regressions, generate insights, and drive cross-functional adoption while mentoring engineers and defining standards for trustworthy AI.
Top Skills: AgentsAi InfrastructureEvaluation SystemsLlmsRetrieval Architectures

What you need to know about the Montreal Tech Scene

With roots dating back to 1642, Montreal is often recognized for its French-inspired architecture and cobblestone streets lined with traditional shops and cafés. But what truly sets the city apart is how it blends its rich tradition with a modern edge, reflected in its evolving skyline and fast-growing tech industry. According to economic promotion agency Montréal International, the city ranks among the top in North America to invest in artificial intelligence, making it le spot idéal for job seekers who want the best of both worlds.

Key Facts About Montreal Tech

  • Number of Tech Workers: 255,000+ (2024, Tourisme Montréal)
  • Major Tech Employers: SAP, Google, Microsoft, Cisco
  • Key Industries: Artificial intelligence, machine learning, cybersecurity, cloud computing, web development
  • Funding Landscape: $1.47 billion in venture capital funding in 2024 (BetaKit)
  • Notable Investors: CIBC Innovation Banking, BDC Capital, Investissement Québec, Fonds de solidarité FTQ
  • Research Centers and Universities: McGill University, Université de Montréal, Concordia University, Mila Quebec, ÉTS Montréal

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account