Inferact logo
AI/ML
DevTools/Cloud

Inferact

AI inference infrastructure startup

CORE INFO

$150M
Total Funding
Seed
Round
2025
Founded
11-50
Team Size
San Francisco, California, USA
Headquarters

Inferact is a San Francisco-based startup that specializes in optimizing AI inference for large language models, offering solutions like the vLLM inference engine to enterprise AI teams and cloud service providers.

  • Inferact secured $150 million in seed funding, led by Andreessen Horowitz and Lightspeed Venture Partners, with a post-money valuation of $800 million.techcrunch.com
  • TechCrunch reported on Inferact's transition from an open-source project to a VC-backed startup, highlighting its significant seed funding and commercialization efforts.techcrunch.com

WHY WE WOULD WORK AT INFERACT

Innovative Technology

Join a team pioneering the vLLM open-source inference engine, designed to optimize large language model performance and accelerate AI progress.

Open Source Commitment

Be part of a company dedicated to advancing the open-source vLLM project, ensuring that your work benefits the broader AI community.

Strong Funding Support

With $150 million in seed funding from top investors like Andreessen Horowitz, Inferact is well-positioned for growth and innovation.

Dynamic Team Culture

Work alongside a small, agile team of experts who are passionate about AI infrastructure and committed to collaboration and innovation.

Impactful Mission

Contribute to a mission focused on making AI inference cheaper and faster, directly influencing the future of AI technology.

Cutting-Edge Solutions

Engage with advanced technologies like PagedAttention and support for over 500 model architectures, shaping the future of AI inference.

MARKET AND TRACTION

TOTAL ADDRESSABLE MARKET

  • Inferact targets enterprise AI engineering teams, cloud service providers, and generative AI application developers.

  • The growing demand for AI inference solutions positions Inferact within a rapidly expanding market as organizations increasingly adopt AI technologies.
  • KEY METRICS

    ✦ KEY METRIC
  • Total funding raised: $150 million in seed funding.

  • Valuation post-funding: $800 million.

  • Team size: Approximately 11-50 employees.
  • SERVICES

  • Inferact offers an Enterprise vLLM Managed Service that simplifies the deployment of AI inference solutions for businesses.

  • The Inference Optimization Engine enhances performance and reduces costs associated with AI model inference.

  • Distributed LLM Orchestration Layer enables efficient management of multiple AI models across various environments.
  • COMPETITIVE ADVANTAGE

  • Inferact's vLLM engine is open-source, allowing for community-driven improvements and widespread adoption.

  • The company utilizes PagedAttention technology for memory optimization, which enhances the efficiency of AI inference processes.
  • MARKET POSITION

  • Founded in 2025, Inferact is positioned as a significant player in the AI infrastructure landscape, focusing on optimizing large language model inference.

  • With substantial backing from prominent investors like Andreessen Horowitz and Lightspeed Venture Partners, Inferact is well-equipped to scale its operations and technology offerings.
  • PRODUCT AND TECH

    vLLM Inference Engine

    The vLLM inference engine is an open-source solution designed to optimize large language model inference, making it faster and more cost-effective for various applications. This engine serves as the backbone of Inferact's offerings, enabling efficient AI processing across diverse environments.

    Enterprise vLLM Managed Service

    This service provides businesses with a fully managed version of the vLLM engine, allowing organizations to leverage advanced AI inference capabilities without the overhead of managing infrastructure. It ensures high availability and performance tailored to enterprise needs.

    Inference Optimization Engine

    Inferact's Inference Optimization Engine enhances the efficiency of AI model execution by applying advanced techniques to reduce latency and resource consumption. This technology is crucial for organizations looking to deploy AI solutions at scale while minimizing costs.

    Distributed LLM Orchestration Layer

    This orchestration layer facilitates the management and coordination of multiple large language models across distributed systems, ensuring seamless integration and scalability. It is essential for organizations that require robust AI solutions capable of handling complex workloads.

    High-Throughput Model Serving APIs

    These APIs enable developers to serve AI models at high throughput, ensuring rapid response times and efficient resource utilization. They are vital for applications that demand real-time AI inference capabilities, such as chatbots and recommendation systems.

    COMPANY CULTURE

    Values

  • Commitment to open-source development

  • Focus on innovation and optimization

  • Dedication to community benefit

  • Emphasis on collaboration and teamwork

  • Integrity in all business practices
  • Operating Principles

  • Prioritize efficiency in AI inference

  • Foster a culture of continuous improvement

  • Encourage experimentation and learning

  • Maintain transparency with stakeholders

  • Support diversity and inclusion in the workplace
  • Benefits

  • Competitive salary and equity options

  • Flexible work hours and remote work opportunities

  • Access to cutting-edge technology and tools

  • Professional development and training programs

  • Health and wellness benefits
  • Learning & Growth

  • Opportunities for mentorship and coaching

  • Regular knowledge-sharing sessions

  • Support for attending industry conferences

  • Encouragement of personal projects and research

  • Focus on skill development in AI and infrastructure
  • Team Cadence

  • Agile project management practices

  • Regular team check-ins and updates

  • Open channels for feedback and communication

  • Collaborative brainstorming sessions

  • Celebrations of team achievements and milestones
  • COMPENSATION SNAPSHOT

    B+
    Comp score vs cohort
    vs. AI/ML · Seed · US
    Based on publicly available compensation ranges from job postings; exact compensation may vary.
    Engineering IC4n = 5
    $200K – $400Kmedian $300K
    Los Angeles / Bay Area
    Engineering IC5n = 5
    $157K – $314Kmedian $236K
    Aggregate
    US-only in v1 · only categories with n > 2 shown · in-person / hybrid / remote treated separately where n ≥ 3