← Back to results

MTS, Inference Performance Visibility

Lead the design and development of performance analysis tools for cutting-edge ML accelerator hardware at Etched.

Location
San Jose, United States
Compensation
Not disclosed
Level
senior
Type
full time · On-site

Posted by employer 8 months ago

First seen on Joblaze 1 day ago

Last verified on the company career page 1 day ago

Apply at Etched → Save job Scanned from etched.com

What you'll build

  • Lead the design and architecture of a performance analysis suite
  • Develop methods to capture performance data from ML accelerator hardware
  • Implement tracing for host-side API calls and system-level events
  • Design techniques to correlate performance events across systems
  • Build analysis modules to interpret collected trace and counter data

Must have

  • Strong proficiency in C++ or Rust
  • Deep understanding of computer architecture
  • Proven experience in low-level performance analysis
  • Experience with performance analysis tools
  • Experience working close to hardware

Nice to have

  • Direct experience developing performance analysis tools
  • Experience with ML accelerator architectures
  • Experience with kernel-mode driver development
  • Understanding of compiler internals
  • In-depth knowledge of the PCIe protocol

Not disclosed in this posting: compensation, years of experience, visa sponsorship.

Benefits

Wellness Benefits Daily Lunch & Dinner Housing Subsidy Unlimited Compute Budget Health Insurance Relocation Assistance

Joblaze summary

In this role, the engineer will lead the development of a performance analysis tool for a custom ML accelerator, focusing on data collection, processing, and visualization to enhance workload understanding. Key skills include proficiency in C++ or Rust, a solid grasp of computer architecture, and experience with performance analysis tools. This position is suited for someone with a strong background in low-level performance profiling and a deep understanding of complex hardware systems. The team operates in a collaborative environment, emphasizing the integration of engineering and research.

Joblaze insights

  • Listed yesterday — first seen on Joblaze September 21, 2026. Last confirmed on Etched's careers page September 21, 2026.
  • Python appears in 53.1% of 544 comparable senior ai/ml roles in United States; PCIe appears in 0.2% of 544 comparable senior ai/ml roles in United States.

Quick facts

Is the MTS, Inference Performance Visibility role remote?
No — this is an on-site role in San Jose, United States.
Where is the role based?
Etched is hiring for this position in San Jose, United States.
What's the tech stack?
Joblaze extracted these technologies from the posting: AMD uprof, C++, ETW, Intel VTune, ML Accelerator, NVIDIA Nsight.
What seniority level is this role?
Etched targets senior candidates for this position.
Is this full-time or contract?
Full-time for this MTS, Inference Performance Visibility role at Etched.

From the original posting

Job Summary

Join our team and take the lead in illuminating the performance landscape of our cutting-edge ML accelerator. We are seeking a highly skilled engineer to design and develop a sophisticated performance analysis tool, tailored specifically for our hardware. You will be instrumental in creating the essential tooling that enables our ML engineers and customers to understand workload behavior, identify performance bottlenecks, and unlock the full potential of our hardware, accelerating the most demanding ML applications in the world. This is a unique opportunity to shape performance analysis for novel hardware from the ground up.

Key responsibilities

  • Lead the design and architecture of a comprehensive performance analysis suite, including data collection mechanisms, data processing pipelines, analysis engines, and user interfaces (CLI and/or GUI).

  • Develop robust methods to capture performance data directly from our custom ML accelerator hardware (e.g., hardware performance counters, execution unit status, memory access patterns) via driver interfaces or other mechanisms.

  • Implement tracing for host-side API calls (runtime libraries, driver interactions) and system-level events (CPU activity, PCIe traffic, memory usage, network contention) related to our workloads.

  • Design and implement techniques to accurately correlate performance events across the host CPU, device driver, PCIe bus, multiple accelerators, and multiple hosts, ensuring precise time synchronization.

  • Build analysis modules to automatically interpret collected trace and counter data, identifying key performance limiters (e.g., compute-bound, memory bandwidth-bound, latency-bound, PCIe-bound, specific hardware bottlenecks).

  • Develop intuitive visualizations (timelines, dependency graphs, resource utilization charts, statistical summaries) to clearly communicate performance characteristics and bottlenecks to users.

  • Work closely with hardware architects, firmware engineers, driver developers, compiler engineers, and ML application engineers to understand their needs, define tool requirements, and provide expert guidance on performance analysis and optimization using the tool.

Representative projects

  • Architect and implement the core data collection framework for hardware performance counters on a custom PCIe-based accelerator.

  • Develop a kernel driver module or user-space service for low-overhead tracing of accelerator activity.

  • Design and build a correlated timeline view visualizing CPU API calls, driver submissions, PCIe transfers, and accelerator execution units.

  • Create an analysis pass to detect and quantify memory access inefficiencies or PCIe bandwidth saturation while transacting on a PCIe-attached accelerator.

You may be a good fit if you have

  • Strong proficiency in C++ or Rust

  • Proficiency in Python is a plus

  • Deep understanding of computer architecture (CPU, GPU, accelerators), memory hierarchies (caches, DRAM), and interconnects (especially PCIe).

  • Proven experience in low-level performance analysis, profiling, and bottleneck identification on complex hardware systems (GPUs, CPUs, FPGAs, or custom accelerators).

  • Experience with performance analysis tools (e.g., NVIDIA Nsight, AMD uProf, Intel VTune, perf, Tracy, ETW).

  • Experience working close to hardware, potentially reading performance counters or interacting directly with device drivers.

Strong candidates may also have experience with (Nice-to-have qualifications)

  • Direct experience developing performance analysis or debugging tools.

  • Experience with ML accelerator architectures (GPUs, TPUs, etc.).

  • Experience with kernel-mode driver development (Linux or Windows).

  • Understanding of compiler internals, code generation, and optimization.

  • In-depth knowledge of the PCIe protocol and analysis tools (PCIe analyzers).

  • Experience with multi-chip or multi-host accelerator systems (e.g., TPU pods, or NVidia DGX clusters)

  • Experience with firmware or embedded systems development.

  • Experience with hardware description languages (Verilog, VHDL) or hardware verification.

Benefits

  • Medical, dental, and vision packages with generous premium coverage

    • $500 per month credit for waiving medical benefits

  • Housing subsidy of $2.5k per month for those living within walking distance of the office

  • Daily lunch + dinner in our office

  • Unlimited compute budget subject to ROI justification

 

Standard company text repeated across Etched's postings is omitted here.

Similar positions

Etched
Head of Inference Performance Visibility
Etched · San Jose, United States
Etched
Performance Tools Intern
Etched · San Jose, United States
Etched
Performance Modeling Engineer
Etched · San Jose, United States
Etched
Applied AI Engineer, Kernel Performance
Etched · San Jose, United States
Etched
Inference Software Engineer
Etched · San Jose, United States