Clockwork Systems, Inc.
AI infrastructure software for resilient GPU clusters
CORE INFO
Clockwork sells software to enterprises, hyperscalers, GPU-cloud providers, and AI infrastructure operators that run large distributed workloads. Its platform monitors GPU clusters, detects infrastructure failures, and reroutes or migrates workloads so training and inference jobs can keep running.
- Latest financing: $31 million announced October 5, 2026, co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures, with NEA and e& Capital participating.prnewswire.com
- LinkedIn uses Clockwork's LinkPass across its AI infrastructure, preventing tens of thousands of GPU-hours of downtime per month, according to the company.prnewswire.com
- TechCrunch reported Clockwork's $21 million Series A in 2022, which supported its work on synchronizing server clocks for distributed computing.techcrunch.com
WHY WE WOULD WORK AT CLOCKWORK SYSTEMS, INC.
Shape AI Infrastructure
Help build the software layer that makes large GPU clusters more observable, efficient, and resilient as AI workloads scale.
Solve Costly Failures
Work on LinkPass and TorchPass to keep training, inference, and reinforcement-learning workloads running through network, GPU, and node failures.
Push Systems Research
Apply Stanford-rooted research in distributed systems, precision clock synchronization, RDMA, and GPU workload migration to hard infrastructure problems.
Join a Scaling Company
Clockwork has raised $73M, including a recent $31M financing, and is expanding FleetIQ across enterprises, hyperscalers, and GPU-cloud providers.
Work With Strong Builders
Collaborate with Stanford researchers, distributed-systems experts, and leaders who helped build companies including VMware, Urban Engines, Nimble Storage, and Sysdig.
See Real-World Impact
Help customers such as LinkedIn, Together AI, WhiteFiber, Uber, and Nebius reduce wasted GPU time and improve production reliability.
MARKET AND TRACTION
GROWTH TACTICS
NOTABLE CUSTOMERS
KEY METRICS
✦ KEY METRICMARKET POSITION
COMPETITIVE ADVANTAGE
PRODUCT AND TECH
FleetIQ AI Fabric
A software-driven AI infrastructure platform combining workload observability, fault tolerance, performance optimization, dynamic traffic control, and cluster-readiness auditing. It helps GPU operators improve utilization, job-completion times, and resilience across distributed AI workloads.
LinkPass Network Failover
A network-resilience technology that detects link failures and reroutes distributed training, reinforcement-learning, and inference traffic across healthy paths. It helps prevent network faults from crashing multi-node workloads or forcing checkpoint restarts.
TorchPass Workload Resilience
A fault-tolerance platform for GPU, node, software, maintenance, and preemption failures. TorchPass supports live GPU migration, multi-node snapshots, application checkpoints, and recovery without restarting entire training jobs.
FleetLens Precision Telemetry
An observability and cluster-audit capability that uses high-precision telemetry to identify the GPU, node, or network link slowing a workload and validate infrastructure before production deployment. Clockwork's Global Clock Sync technology provides near-nanosecond synchronization across hosts, switches, and NICs for precise distributed-system measurements.