Devoted Studios Logo

Devoted Studios

Performance Engineer

Reposted 12 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in Montréal, QC, CAN
Senior level
In-Office or Remote
Hiring Remotely in Montréal, QC, CAN
Senior level
Lead performance initiatives across CPU, GPU, memory, networking, and profiling tooling in Unreal Engine 5. Profile, diagnose, and implement C++ fixes, build automated regression tests and profiling tools, and collaborate with cross-disciplinary teams to ensure platform memory and performance budgets are met.
The summary above was generated by AI

Devoted Studios is a globally remote game development company specializing in Co-development, Porting, and End-to-End Art Production for the global gaming industry. We collaborate across time zones to support projects on all major platforms, engines and styles - from AAA titles to emerging technologies.

Our team includes world-class talents who bring deep expertise in external development, pipeline optimization, and creative problem-solving. Whether it’s porting games to new systems, enhancing gameplay features, or crafting stunning visuals, Devoted Studios operates as a trusted, flexible extension of our partners’ internal teams.

We are proud to be the development partner of choice for industry leaders such as:

    • Epic
    • Xbox
    • Meta
    • Obsidian Entertainment
    • Riot
    • Gearbox Software

In this role, you will own and drive the development of core systems and performance initiatives from planning and design through implementation and optimization. Working closely with engineers and cross-disciplinary teams, you will provide technical leadership and feature ownership.

CPU & Game Thread Optimisation (~25%)

• Profile and optimise the game thread, render thread, and RHI thread in Unreal Engine — identifies cross-thread bottlenecks, tick overhead, and task graph inefficiency

• Diagnose and resolve hitching: garbage collection pauses, streaming stalls, async loading conflicts, physics simulation spikes

• Reduce tick overhead: identifies UObjects and ActorComponents with unnecessary tick enabled, unnecessary work in hot paths, incorrect tick group assignments

• Optimise gameplay and engine code at the C++ level — not just profiling and handing recommendations to another engineer; writes and ships the fix

• Instruments custom profiling markers (SCOPE_CYCLE_COUNTER, CSV_SCOPED_TIMING_STAT, custom Unreal Insights channels) for ongoing regression detection

GPU & Rendering Pipeline Optimisation (~25%)

• Profiles rendering performance using RenderDoc, PIX, NSight, Xcode GPU Frame Capture, and Unreal's GPU Visualiser — isolates draw call overhead, overdraw, shader complexity, and GPU memory bandwidth

• Works with the rendering team to optimise material pipelines: shader permutation reduction, LOD tuning, occlusion culling, distance field shadow optimisation

• Identifies and resolves GPU memory pressure: texture streaming configuration, render target pooling, VRAM budget management per platform

• Understands the rendering architecture deeply enough to distinguish CPU-bound vs GPU-bound vs bandwidth-bound frames and route the fix correctly

• Targets platform-specific GPU constraints: the PS5/Xbox Series GPU architecture differs meaningfully from PC; mobile (Adreno, Mali, Apple GPU) has its own thermal and bandwidth profile

Memory Profiling & Leak Detection (~20%)

• Owns memory budget compliance on client engagements: tracks resident set size, texture memory, audio memory, and engine overhead against platform limits

• Finds and eliminates memory leaks in C++ using both automated tooling (ASan, Valgrind on non-console platforms, custom allocator hooks) and Unreal's LLM (Low-Level Memory tracker)

• Manages memory fragmentation: understands pooled allocators, custom memory arenas, and when Unreal's default allocation strategy creates fragmentation under long session times

• Console memory work is the hardest version of this: fixed RAM ceilings with no swap. Has shipped inside those constraints.

• Builds memory regression tests that run in CI — not just point-in-time snapshots

Networking / Latency Diagnostics (~10%)

• Profiles and optimises Unreal's replication system: identifies bandwidth-heavy actors and components, tunes replication frequency, prioritises replication correctly under load

• Diagnoses client-server latency: distinguishes network latency from server frame time from client simulation cost — does not conflate them

• Familiar with NetStat, Unreal's network profiler, and packet-level capture when needed

• Scope: supports multiplayer engagements where networking is a performance constraint — not the primary discipline on pure-networking architecture

Profiling Tool Development & Automation (~20%)

• Builds internal profiling tools in C++ (UE plugins) and Python — automated perf validation scripts, CI-integrated regression detection, custom Unreal Insights data channels

• Writes automated performance regression tests: baseline capture, threshold alerting, trend visualisation — so regressions are caught in the build, not in the client review

• Builds or improves developer-facing profiling UX: makes it easier for other engineers on the engagement to find their own bottlenecks without specialist involvement

• Documents profiling methodology and tool usage: other engineers can reproduce the analysis without a knowledge transfer session


Requirements

• 5+ years of professional game engineering experience with at least 3 years in a dedicated performance, engine, or low-level systems role

• At least 2 shipped commercial titles where performance was a named ownership area — not a general engineering role that included some profiling

• C++ at a deep level: understands memory layout, cache efficiency, branch prediction, SIMD, and how the compiler's optimisation decisions affect runtime behaviour

• Unreal Engine 5 profiling expertise: fluent in Unreal Insights, Stat commands, LLM, GPU Visualiser — not just aware of them, but has used them to find real problems on real projects

• Cross-platform experience: has shipped on at least two platform types (PC + console, or console + mobile) and understands the meaningfully different constraint profiles

• Python proficiency: writes automation scripts, CI integration, data processing for perf telemetry — not a Python developer, but not blocked by Python tasks

• Structural fix mindset: profiles, identifies root cause, writes the C++ fix. Does not stop at the recommendation.

• Can explain a performance problem to a gameplay engineer or producer without requiring them to understand the GPU pipeline


Preferred

• Console certification performance submissions: has prepared and passed Sony/Microsoft performance gates

• Mobile performance work on Android (Adreno/Mali/Dimensity) and/or iOS (Apple Silicon GPU) — thermal throttling, tiling architecture, bandwidth constraints

• Custom allocator or memory pool implementation in C++

• Shader authoring knowledge: HLSL/GLSL fluency sufficient to read a shader and identify unnecessary ALU or texture sample cost

• Familiarity with Tracy, Optick, or custom in-house profiling frameworks beyond Unreal's built-in tooling

• Experience optimising Nanite, Lumen, or other UE5-specific rendering features at engine depth



Benefits

This is a full-time employment role with standard benefits in Quebec province.

Similar Jobs

3 Days Ago
Remote
Canada
Senior level
Senior level
Cloud • Security • Software • Generative AI
Develop and optimize Elasticsearch performance across distributed and serverless architectures. Build performance models, profiling and benchmarking workflows, regression detection, and AI-assisted optimization harnesses. Analyze bottlenecks in logging, metrics, vector search, and ES|QL; implement scalable, thread-safe, high-performance Java code; and collaborate across teams through technical reviews and guidance. Experience with JVM internals, distributed systems, performance engineering, and benchmarking is required.
Top Skills: C++ElasticsearchEs|QlJavaJmhJvmRallyRustSimdVector Search
3 Days Ago
Remote
Canada
Expert/Leader
Expert/Leader
Cloud • Security • Software • Generative AI
Lead architectural and code-level performance engineering for Elasticsearch. Own major optimization initiatives, develop performance models, profile distributed systems, identify bottlenecks, build benchmarks and regression detection, and improve scalability across stateful and serverless architectures. Develop AI-assisted optimization and benchmarking harnesses, collaborate across teams, and mentor engineers. Work includes Java and JVM optimization, concurrency, distributed systems, vector search, storage engines, and native library integration.
Top Skills: C++Distributed SystemsElasticsearchEs|QlFlamegraphsJavaJmhJvmRallyRustSimdVector Search
15 Days Ago
Remote
Canada
Mid level
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Drive end-to-end ML model inference performance: build kernel- and system-level performance models, optimize kernel microcode and compiler algorithms, debug runtime performance on system and cluster, and develop tooling to visualize and analyze performance data from the Wafer Scale Engine and compute cluster.
Top Skills: C++Cerebras Wafer Scale EngineCompilersCpu/Gpu SimulatorsHpcKernel MicrocodePerformance Profiling ToolsPython

What you need to know about the Montreal Tech Scene

With roots dating back to 1642, Montreal is often recognized for its French-inspired architecture and cobblestone streets lined with traditional shops and cafés. But what truly sets the city apart is how it blends its rich tradition with a modern edge, reflected in its evolving skyline and fast-growing tech industry. According to economic promotion agency Montréal International, the city ranks among the top in North America to invest in artificial intelligence, making it le spot idéal for job seekers who want the best of both worlds.

Key Facts About Montreal Tech

  • Number of Tech Workers: 255,000+ (2024, Tourisme Montréal)
  • Major Tech Employers: SAP, Google, Microsoft, Cisco
  • Key Industries: Artificial intelligence, machine learning, cybersecurity, cloud computing, web development
  • Funding Landscape: $1.47 billion in venture capital funding in 2024 (BetaKit)
  • Notable Investors: CIBC Innovation Banking, BDC Capital, Investissement Québec, Fonds de solidarité FTQ
  • Research Centers and Universities: McGill University, Université de Montréal, Concordia University, Mila Quebec, ÉTS Montréal

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account