AI/ML
DevTools/Cloud
Enterprise SaaS

Clockwork Systems, Inc.

AI infrastructure software for resilient GPU clusters

CORE INFO

$73M
Total Funding
2018
Founded
Palo Alto, CA
Headquarters

Clockwork sells software to enterprises, hyperscalers, GPU-cloud providers, and AI infrastructure operators that run large distributed workloads. Its platform monitors GPU clusters, detects infrastructure failures, and reroutes or migrates workloads so training and inference jobs can keep running.

  • Latest financing: $31 million announced October 5, 2026, co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures, with NEA and e& Capital participating.prnewswire.com
  • LinkedIn uses Clockwork's LinkPass across its AI infrastructure, preventing tens of thousands of GPU-hours of downtime per month, according to the company.prnewswire.com
  • TechCrunch reported Clockwork's $21 million Series A in 2022, which supported its work on synchronizing server clocks for distributed computing.techcrunch.com

WHY WE WOULD WORK AT CLOCKWORK SYSTEMS, INC.

Shape AI Infrastructure

Help build the software layer that makes large GPU clusters more observable, efficient, and resilient as AI workloads scale.

Solve Costly Failures

Work on LinkPass and TorchPass to keep training, inference, and reinforcement-learning workloads running through network, GPU, and node failures.

Push Systems Research

Apply Stanford-rooted research in distributed systems, precision clock synchronization, RDMA, and GPU workload migration to hard infrastructure problems.

Join a Scaling Company

Clockwork has raised $73M, including a recent $31M financing, and is expanding FleetIQ across enterprises, hyperscalers, and GPU-cloud providers.

Work With Strong Builders

Collaborate with Stanford researchers, distributed-systems experts, and leaders who helped build companies including VMware, Urban Engines, Nimble Storage, and Sysdig.

See Real-World Impact

Help customers such as LinkedIn, Together AI, WhiteFiber, Uber, and Nebius reduce wasted GPU time and improve production reliability.

MARKET AND TRACTION

GROWTH TACTICS

  • Expanding fault-tolerance products across AI training, inference, and reinforcement learning, while increasing enterprise adoption.

  • Scaling distribution through cloud partners and GPU-cloud providers, including commercial relationships with Together AI and WhiteFiber.
  • NOTABLE CUSTOMERS

  • Named customers and adopters include LinkedIn, Together AI, WhiteFiber, Uber, Wells Fargo, Nebius, Nscale, and DCAI.

  • LinkedIn has deployed LinkPass across its AI infrastructure, while Together AI is bringing TorchPass to market through its GPU-cluster services.
  • KEY METRICS

    ✦ KEY METRIC
  • Clockwork has raised $73 million in total disclosed funding, including a $31 million financing announced in October 2026.

  • The company reports that FleetIQ can improve GPU-cluster utilization and job-completion times by 1.1–1.5x and reduce disruptive failures by more than 90%.

  • LinkedIn reportedly prevents tens of thousands of GPU-hours of downtime per month using Clockwork technology.
  • MARKET POSITION

  • Clockwork serves enterprises, hyperscalers, GPU-cloud providers, neoclouds, and AI infrastructure operators running large distributed GPU environments.

  • Its current commercial positioning centers on AI infrastructure software for training, inference, reinforcement learning, and high-performance computing.
  • COMPETITIVE ADVANTAGE

  • FleetIQ provides a software control and resilience layer spanning workload observability, network fault detection and rerouting, GPU and node fault tolerance, live GPU migration, snapshots, traffic control, and precision clock synchronization.

  • The platform supports Ethernet, InfiniBand, RoCE, NVIDIA and AMD GPUs, custom accelerators, NCCL, RCCL, Kubernetes, Slurm, and major AI workload frameworks.
  • PRODUCT AND TECH

    FleetIQ AI Fabric

    A software-driven AI infrastructure platform combining workload observability, fault tolerance, performance optimization, dynamic traffic control, and cluster-readiness auditing. It helps GPU operators improve utilization, job-completion times, and resilience across distributed AI workloads.

    LinkPass Network Failover

    A network-resilience technology that detects link failures and reroutes distributed training, reinforcement-learning, and inference traffic across healthy paths. It helps prevent network faults from crashing multi-node workloads or forcing checkpoint restarts.

    TorchPass Workload Resilience

    A fault-tolerance platform for GPU, node, software, maintenance, and preemption failures. TorchPass supports live GPU migration, multi-node snapshots, application checkpoints, and recovery without restarting entire training jobs.

    FleetLens Precision Telemetry

    An observability and cluster-audit capability that uses high-precision telemetry to identify the GPU, node, or network link slowing a workload and validate infrastructure before production deployment. Clockwork's Global Clock Sync technology provides near-nanosecond synchronization across hosts, switches, and NICs for precise distributed-system measurements.

    COMPANY CULTURE

    Values

  • Research-driven innovation

  • Solve fundamental technical problems

  • Build infrastructure that makes AI more reliable and efficient

  • Carry others with us as the company grows
  • Operating Principles

  • Prioritize customer outcomes and production reliability

  • Turn complex infrastructure challenges into practical software solutions

  • Make communication fabrics observable, adaptive, and resilient

  • Celebrate meaningful technical progress
  • Work Style

  • Collaborative, technically rigorous, and mission-driven

  • Small-team environment where talented people can do their best work

  • Focused on high-impact challenges in AI and distributed systems

  • Connect research foundations with real-world customer deployments
  • Learning & Growth

  • Work on fundamental challenges in AI performance and GPU infrastructure

  • Learn from Stanford-rooted research and deep distributed-systems expertise

  • Grow alongside an emerging software-driven AI infrastructure category
  • Benefits

  • Free lunch

  • Free gym

  • TGIF events

  • Company benefits