Mediabistro logo
job logo

Senior / Staff ML Compiler Engineer

Crescent City Recruitment Group · San Jose, CA, USA ·

Job type:
Full Time

Senior / Staff ML Compiler Engineer

Location: San Jose or Irvine, CA | Full-Time

About Our Client

Our client is a fast-growing fabless semiconductor company building next-generation, energy-efficient domain-specific processors designed for edge AI, wireless communications, advanced radar, computer vision, and autonomous systems.

Founded by industry veterans, our client is developing breakthrough System-on-Chip (SoC) architectures leveraging RISC-V and advanced compute technologies to deliver real-time intelligence at the sensor edge.

This is an opportunity to join a highly innovative engineering team working at the intersection of semiconductor architecture, AI acceleration, and compiler technology.

Position Summary

We are seeking a

Senior / Staff ML Compiler Engineer

to develop and optimize compiler technologies that unlock the performance of next-generation custom silicon.

This is

not a traditional application software engineering role.

The ideal candidate brings deep expertise in compiler backend development, code generation, runtime optimization, hardware/software co-design, and performance tuning for compute-intensive architectures.

This individual will work closely with architecture, silicon, systems, and AI teams to ensure software fully enables our client's processor roadmap.

Key Responsibilities

Compiler Architecture & Development
Design, develop, and optimize compiler infrastructure for custom processor and accelerator architectures
Build compiler backends, optimization passes, code generation workflows, and execution pipelines
Improve instruction scheduling, register allocation, memory access efficiency, and execution performance
Develop graph-level and operator-level optimization strategies for AI workloads
Support compiler enablement for emerging compute architectures
Runtime & Performance Optimization

Develop runtime systems and execution frameworks for optimized inference and compute performance
Profile workloads and identify bottlenecks impacting latency, throughput, memory bandwidth, and power efficiency
Debug performance issues across simulation, emulation, and silicon environments
Build internal benchmarking and performance analysis tools
Hardware / Software Co-Design

Partner with architecture and hardware teams on next-generation processor development
Analyze workload behavior to influence architecture decisions
Help drive tradeoff analysis involving performance, memory efficiency, latency, and power consumption
Collaborate with silicon teams during bring-up and optimization cycles
AI / ML Workload Enablement

Optimize execution of machine learning and inference workloads on custom accelerators
Support deployment of workloads involving:
computer vision
edge AI
radar
sensing
autonomous systems
signal processing
Collaborate with AI software teams on model optimization and framework integration
Required Qualifications

BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or related discipline
7+ years of relevant experience (Senior level)
10+ years of relevant experience (Staff level)
Strong C/C++ programming expertise
Deep experience with compiler development and optimization
Hands-on expertise with one or more of:
LLVM
MLIR
TVM
XLA
GCC
custom compiler frameworks
Experience in:
compiler backend development
code generation
optimization passes
runtime systems
performance profiling
Strong understanding of:
computer architecture
memory systems
instruction scheduling
register allocation
parallel execution
low-level performance optimization
Preferred Qualifications

Semiconductor industry experience
Experience with AI accelerators, NPUs, DSPs, GPUs, or custom compute architectures
Knowledge of RISC-V architectures
Experience with:
graph optimization
quantization
inference optimization
edge AI deployment
embedded systems
HW/SW co-design
Familiarity with autonomous systems, radar, signal processing, or computer vision workloads