Etched has raised a total of $800 million across four previously unannounced funding rounds, emerging from stealth with a working A0 silicon chip (Sohu), over $1 billion in signed customer contracts for its inference systems, and plans to ship first racks in summer 2026.
What is Etched?
The company, founded in 2022 by Harvard dropouts Gavin Uberti (CEO) and Chris Zhu, along with Robert Wachen (President), specializes in custom ASICs and rack-scale systems optimized exclusively for transformer based AI inference workloads. This narrow focus (targeting LLMs, sparse MoEs, long context, and agentic applications) distinguishes it from general purpose GPUs. Key leadership includes experienced executives like CTO Mark Ross (ex Cypress Semiconductor), VP Platform Brian Loiler (ex NVIDIA, HGX/DGX systems), and others from Broadcom, Intel, TSMC, SK Hynix, Google TPUs, and DeepMind.
The $800M total includes a major $500 million round closed around December 2025/January 2026, led by Stripes with participation from Peter Thiel, Positive Sum, Ribbit Capital, Jane Street, Hudson River Trading, Two Sigma, and others, at a ~$5 billion post money valuation. Jane Street has invested over $100 million across rounds. Additional backers encompass VentureTech Alliance (strategic tie to TSMC), AI luminaries like Geoffrey Hinton, Andrej Karpathy, Fei-Fei Li, Arthur Mensch (Mistral), Aidan Gomez (Cohere), Noam Brown (OpenAI), and others such as Stanley Druckenmiller. Earlier rounds brought the cumulative to $800M, reflecting strong conviction from trading firms, hedge funds, and domain experts.
This capital supports vertical integration: chip design, packaging, racks, software, a Taiwan factory, San Jose data center/test lab, and rapid scaling toward gigawatt scale deployment.

What is Etched’s technology?
Etched’s approach embeds transformer architecture directly into silicon for superior efficiency on its target workloads, bypassing the overhead of general purpose hardware. Key breakthroughs include:
- Low Voltage Inference (LVI): Runs math blocks at under half the voltage of typical AI chips, enabling multiple times higher FLOPs density and sustained 80%+ peak FLOPs utilization on trillion parameter sparse MoEs without thermal throttling. This addresses power and heat limitations in high throughput scenarios.
- Cluster Scale Memory (CSM): A hybrid HBM/SRAM design with proprietary ultra low latency, high bandwidth interconnects for a shared memory pool. This improves decode latency and interactivity while balancing capacity, avoiding SRAM only or optics tradeoffs.
The Sohu chip (fabricated on TSMC N4P) powers rack-scale “frontier inference clusters” optimized for both prefill and decode. Early claims (from prior disclosures) positioned an 8-chip Sohu server at ~500,000 tokens/second on Llama 70B, potentially equating to 10-20x throughput advantages over equivalent H100 or B200 setups in transformer inference, with better cost and power metrics. The company reports state of the art (SOTA) results in customer tests for throughput, latency, and efficiency.
Production ramp is underway, with A0 silicon validated and racks in testing. Co-design with AI companies, cloud providers, and hyperscalers has driven these outcomes.
Etched targets the exploding inference market, where cost, power, and speed dominate as models scale. By specializing in transformers (the dominant architecture for frontier models), it achieves higher utilization and density than GPUs, which must support diverse workloads. This could deliver significant TCO savings for hyperscalers and enterprises running massive inference.
Traction highlights include over $1B in customer contracts for upcoming racks, signaling strong pre validation despite the unannounced status until now. First shipments are imminent, positioning Etched for revenue in 2026. The deep TSMC partnership via VentureTech and manufacturing expertise (e.g., iPhone/MacBook ramps) de-risks scaling.

Recommended: Runpod Raises $100 Million In Funding Led By Summit Partners
In the AI chip race, Etched challenges NVIDIA’s dominance in inference while competing with other specialists (e.g., Groq for low latency, Cerebras, SambaNova, or hyperscaler custom silicon like Google TPUs/AWS Inferentia). NVIDIA retains advantages in ecosystem, software (CUDA), and versatility, but inference specific ASICs like Sohu can offer superior performance per dollar and power efficiency for high volume transformer serving. Success hinges on execution, software maturity, and proving claims at scale amid supply chain and yield challenges common to new ASICs.
Strengths include elite talent, substantial capital, proven manufacturing pedigree, early revenue commitments, and a focused bet on transformers amid AI’s inference boom. Vertical integration and co-design accelerate iteration toward gigawatt clusters.
Challenges involve ASIC development risks (though A0 success mitigates this), competition from NVIDIA’s ongoing innovations (e.g., Blackwell), the need for robust software/ecosystem support, and geopolitical/manufacturing dependencies. Transformer dominance is a reasonable but not guaranteed assumption long term.
Etched’s $800M raise and $1B+ pipeline mark it as a serious contender in specialized AI hardware. With racks shipping soon and strong backing, it is poised for rapid commercialization, potentially capturing significant share in the inference market if performance and scaling targets are met. The company is hiring aggressively to support expansion.
Please email us your feedback and news tips at hello(at)techcompanynews.com

