apposters.com

Inferact Secures $150M to Revolutionize AI Inference

January 27, 2026, 9:38 pm
a16z
a16z
FinTechPlatformDataHealthTechTechnologySoftwareServiceITCloudGaming
Employees: 51-200
Lightspeed Venture Partners
Lightspeed Venture Partners
PlatformDataFinTechTechnologyServiceSoftwareCloudProductBusinessIT
Location: United States, California, Menlo Park
Employees: 51-200
Founded date: 2000
Sequoia Capital
Sequoia Capital
DataPlatformSoftwareFintechTechnologyServiceITCloudSecurityHealthTech
Location: United States, California, Menlo Park
Employees: 51-200
Founded date: 1972
University of California, Berkeley
University of California, Berkeley
CollegeContentEdTechHomeITLearnPagePublicResearchUniversity
Location: United States, California, Berkeley
Employees: 10001+
Founded date: 2015
Total raised: $250K
Inferact, a San Francisco-based AI firm, just secured a significant $150 million in seed funding. This substantial investment, co-led by industry giants Andreessen Horowitz and Lightspeed Venture Partners, signals strong confidence in the company's groundbreaking technology. Inferact develops a high-performance AI inference engine. Its flagship, open-source vLLM project, powers large language models (LLMs) globally. A core innovation, PagedAttention, dramatically improves efficiency. This advanced memory management technique reduces operational costs and enhances performance for LLM deployment. The funds will fuel team expansion and vLLM's commercialization. This strategic move positions Inferact as a critical player in the evolving AI infrastructure landscape, simplifying complex AI model deployment for enterprises worldwide and transforming how businesses scale their AI capabilities. The firm addresses a crucial need for faster, more cost-effective AI operations.

San Francisco, CA — Inferact, a new force in artificial intelligence infrastructure, has closed a massive $150 million seed funding round. This significant capital injection underscores investor belief in the company's innovative approach to AI inference. The round saw co-leadership from prominent venture capital firms Andreessen Horowitz and Lightspeed Venture Partners. Other key participants included Sequoia Capital, Altimeter Capital, Redpoint Ventures, and ZhenFund. The infusion of funds will accelerate Inferact's mission. It will drive team expansion and finalize the commercial version of its core vLLM technology.

Inferact is poised to transform how businesses deploy and manage large language models (LLMs). The company specializes in a high-performance AI inference engine. Its open-source vLLM project forms the foundation of this engine. This technology is not just about speed. It fundamentally reduces the operational costs associated with running complex LLMs. It simultaneously boosts performance through advanced memory management techniques.

The heart of Inferact's innovation lies in PagedAttention. This proprietary memory management system reimagines GPU memory utilization. PagedAttention operates on principles similar to virtual memory in operating systems. It drastically cuts down memory waste. Typically, LLM inference can see 60-80% memory loss. PagedAttention minimizes this to nearly zero. This efficiency allows companies to process more requests on the same hardware. It maximizes resource utilization. This directly translates to substantial cost savings and enhanced processing power for AI applications.

The vLLM project originated from the UC Berkeley Sky Computing Lab. It has already established a robust ecosystem. It supports over 500 model architectures. It runs efficiently on more than 200 accelerator types. This widespread compatibility enables vLLM to power AI inference at a global scale. The open-source community has contributed significantly to its growth. Over 2,000 contributors actively participate in its development. This collective effort highlights its industry relevance and broad adoption. The code is also integrated within the PyTorch ecosystem, further cementing its position.

Major players already leverage vLLM’s capabilities. Amazon's assistant, Rufus, serving 250 million users, relies on this engine. Roblox's Assistant, processing over a billion tokens weekly, also uses vLLM. LinkedIn’s Hiring Assistant benefits from its efficiency. These high-profile adoptions demonstrate vLLM's robust performance under demanding real-world conditions. Inferact's CEO, Simon Mo, notes that Amazon Web Services and the Amazon Shopping application are also significant users. This widespread enterprise adoption validates Inferact's technology as a critical enabler for modern AI services.

The $150 million seed funding round values Inferact at an impressive $800 million. This valuation reflects investor confidence in the company's market potential. It signals the growing importance of AI infrastructure. The capital will allow Inferact to aggressively expand its engineering and product teams. It will also accelerate the development of its commercial vLLM engine. The goal is to make AI model deployment as straightforward as possible.

Inferact’s leadership team brings deep expertise to the venture. Simon Mo, a co-creator of the vLLM project, serves as CEO. Woosuk Kwon is the original architect of vLLM. He authored the foundational dissertation on the technology. Ion Stoica is also a co-founder. Stoica is a renowned Berkeley professor. He co-founded industry giants Databricks and Anyscale. Their collective vision guides Inferact's strategic direction. They aim to simplify complex AI operations for the enterprise.

The company envisions a future where deploying sophisticated AI models requires minimal effort. Today, extensive infrastructure teams are often necessary for such tasks. Inferact aims to absorb this complexity. Its commercial offerings will streamline the process. The objective is to make advanced AI accessible and manageable for all businesses. This vision suggests a significant shift in AI operational paradigms.

The venture capital world is keenly focused on AI infrastructure. Inferact's funding round is not an isolated event. Another project from Professor Stoica's UC Berkeley lab, SGLang, recently spun out into a startup called RadixArk. RadixArk secured a $400 million valuation just a week prior. These back-to-back investments from a single laboratory send a clear message. Investors see AI infrastructure, particularly deployment and inference, as the next major battleground in the AI revolution.

Inferact stands at the forefront of this critical shift. Its vLLM technology, powered by PagedAttention, offers a compelling solution. It addresses the escalating demand for efficient, cost-effective LLM deployment. The company's strategy combines open-source leadership with commercial innovation. This positions Inferact to become an indispensable partner for enterprises scaling their AI capabilities. Its impact on the future of AI will be substantial. Inferact is building the foundational technology for the next generation of intelligent applications.