Member of Technical Staff
Wafer · San Francisco
No salary listed. Estimated from 259 salary-disclosed Research roles on this board: $235K–$325K (interquartile range; estimate, not the employer's figure).
Apply at WaferFind your warm intro on LinkedInOur mission at Wafer is to maximize intelligence per watt by using AI to optimize AI infrastructure, achieving orders of magnitude better energy and cost efficiency per token.
We believe cheap intelligence is the most essential piece of technology for a future of abundance. We care about building a future where intelligence is "too cheap to meter."
Wafer commercializes these efforts by serving serverless and dedicated inference for open source LLMs at the best performance per dollar. Our core bet is doing this through autonomous optimization of heterogeneous hardware.
What you'll doShip day-zero support for new open-source models, tuned for latency and throughput
Optimize the serving stack: batching, KV cache, speculative decoding, quantization
Write and tune kernels in CUDA, HIP, and Triton for NVIDIA, AMD, TPU, Trainium, D-Matrix, and more.
Design, deploy, and operate heterogeneous clusters across vendors
Run production inference across a mixed fleet: reliability, observability, and cost per token at scale
We score every candidate on seven values:
Infinitely Resourceful
Exceptionalism
Unreasonable Standards
Company Over Self
High EQ
Learns Quickly
First Principles Thinker
$200K base salary + 0.5-1% equity.
Fully covered medical, dental, and vision insurance.
Daily lunch and dinner, unlimited PTO, and parental leave.
$1K/month housing stipend (post-tax) if you live within walking distance (0.5 miles) from the office.
Covered Uber/Waymo from/to office.
Visa sponsorship available.
On-site in San Francisco, five days a week. Small team with massive surface area and ownership. You operate with complete autonomy of how to solve problems, and work with the team to set the direction of your work. We don't see engineers as code writers, but as problem solvers. You will do everything from talking to customers to writing custom GPU kernels in esoteric hardware.