Callosum is positioning as a pre seed horizontal AI infrastructure play, building foundational capabilities around micro-model meshes.
As agentic architectures emerge as the dominant build pattern, Callosum is positioned to benefit from enterprise demand for autonomous workflow solutions. The timing aligns with broader market readiness for AI systems that can execute multi-step tasks without human intervention.
Callosum is an intelligent systems company that develops systems-level software that balances AI workloads across a diverse mix of hardware.
A vertically integrated, topology-aware stack that co-evolves models, workflows, kernels and multiple disparate hardware paradigms—combined with bespoke low-level kernels (including on-die masking), execution-graph-aware caching/prefetching, and cross‑vendor orchestration that together produce orders-of-magnitude gains in cost, latency and capability for heterogeneous, agentic workloads.
They partition workflows into many specialized models (different sizes and capabilities) and route sub-tasks to the smallest/best model for each step. Includes ensemble inferencing (multiple candidates from small models) and automatic discovery/routing across model-hardware combinations to trade accuracy, cost and latency.
Cost-effective AI deployment for mid-market. Creates opportunity for specialized model providers.
End-to-end orchestration of autonomous agents that call tools, interact with environments (web, APIs), maintain trajectory/memory, and coordinate multi-step plans. The infra explicitly optimizes tool interfaces, verification, and multi-agent coordination.
Full workflow automation across legal, finance, and operations. Creates new category of "AI employees" that handle complex multi-step tasks.
Workflows include explicit retrieval stages (selecting and chunking relevant context), long-range retrieval and verification steps integrated into recursive/multi-step model pipelines to augment generation with external context.
Accelerates enterprise AI adoption by providing audit trails and source attribution.
They use secondary/smaller models and ensemble strategies as verifier/validator layers (for coordinates, structured outputs, semantic correctness), plus automatic detection of failure points and rerouting to verification models to enforce safety/correctness.
Accelerates AI deployment in compliance-heavy industries. Creates new category of AI safety tooling.
Callosum builds on Claude Opus 4.5, GPT-5.2, GPT-5, leveraging OpenAI and Anthropic infrastructure with vLLM, SGLang in the stack. The technical approach emphasizes rag.
Multi-agent, multi-model orchestration with recursive model calls and agent handoffs. Structured tool-calls (JSON) are enforced on-silicon; agents maintain trajectory memory and visual working memory; system composes diverse models (vision-language-action, LLMs) and chips into a single workflow executed by a topology-aware runtime.
Execution-graph-aware routing: workflows are decomposed into stages (planning, expansion, retrieval, verification, tool calling) and each stage is routed to the model and chip that best match its latency, cost, and memory profile. The runtime also routes based on failure modes (e.g., GPT-5.2 coordinate errors routed to small 8B verifier), user-specified optimization objectives (cost vs latency vs quality), and discovered Pareto-optimal configurations from benchmarks.
Hands-on hardware-software co-design, neuroscience-inspired AI systems; experience spanning chip design, kernel development, cluster operations, cloud infrastructure, and new model architectures; education/work across Cambridge, Oxford, MIT, Imperial College London.
Previously: Microsoft, ETH Zurich, Intel
Founders' backgrounds in hardware-software co-design, AI systems, and cross-disciplinary expertise align well with Callosum's heterogenous compute vision; strong fit for building intelligent systems that co-evolve with hardware.
partnership led
Target: enterprise
custom
field sales
• Coworker AI as an enterprise partner case study
• Production deployments and feedback from enterprise-oriented use cases
Orchestrating heterogeneous models and hardware to solve multi-agent, multi-modal AI tasks across dynamic, real-world environments
Recursive model invocation is common, but deliberately partitioning recursion stages across heterogeneous models and hardware (co-optimised end-to-end) is a less-explored pattern that unlocks new points on the cost-latency-accuracy Pareto frontier.
Combining cross-cloud heterogeneous scheduling with an execution-graph-aware runtime managing KV caches across chips is a sophisticated integration rarely seen in production LLM stacks.
Moving grammar enforcement fully onto accelerator silicon and achieving O(1) scaling vs. CPU O(B) is a substantial systems innovation that both improves safety (structural validity) and enables expensive downstream strategies (ensembles).
Callosum operates in a competitive landscape that includes NVIDIA, AWS (including Inferentia / Trainium / SageMaker), Cerebras Systems.
Differentiation: Callosum focuses on orchestrating heterogeneous stacks across many chip types and co-evolving models/kernels with varied silicon; NVIDIA focuses on optimizing homogeneous GPU-based stacks and end-to-end ecosystems (hardware + CUDA software). Callosum emphasizes cross-vendor orchestration, topology-aware runtimes, and on-die / specialized kernels on non‑NVIDIA silicon rather than relying primarily on CUDA/GPU homogeneity.
Differentiation: Callosum is cloud-agnostic and orchestrates workloads across multiple cloud providers and next‑gen compute vendors; it also builds custom on-die kernels (e.g., Inferentia2 integration) and a topology-aware runtime to co‑design models and silicon—whereas AWS primarily exposes and operates its own hardware and services and tends toward vertically integrated hyperscaler offerings.
Differentiation: Callosum uses accelerators like Cerebras as one substrate among many and differentiates by orchestrating multiple disparate physical paradigms jointly (photonic, biological, superconducting, conventional ASICs) and optimizing workflows across them. Cerebras is primarily a hardware + software vendor for its own architecture, not an orchestrator across heterogeneous silicon.
Topology-aware KV cache management that uses the workflow execution graph to approximate Bélády’s optimal eviction (evicting nodes furthest from future use), combined with prefetching and hierarchical multi-tier caching across heterogeneous chips — not just an LRU/LFU tweak but a graph-driven runtime that unifies caching across models, context lengths and hardware.
Heterogeneous recursion: decomposing recursive language-model pipelines across different model families and physical substrates (different chips per recursion depth/role) and automatically discovering cost/latency/accuracy tradeoffs. They treat recursion as an allocation/search problem across model+chip pairs rather than scaling one model deeper.
On-die grammar enforcement for structured tool-calls: compiling JSON schemas into finite-state machines and running constrained decoding inside Inferentia2 NeuronCore SBUF (mask in on-chip SRAM). This converts a PCIe roundtrip CPU bottleneck into O(1) on-accelerator masking (microsecond-level cost) and enables cheap ensemble inference at scale.
Per-action heterogeneous model selection in interactive/active-perception agent loops (the 'zoom-step' pattern): at the granularity of individual actions, route verification/localisation tasks to tiny 8B models and planning/global steps to large VLMs — demonstrably improving reliability while massively reducing cost/latency per interaction.
Cross-cloud, cross-instance orchestration with claimed 'cross-vendor GPU networking breakthroughs' to overcome cross-instance networking as the bottleneck. The orchestration is multi-endpoint by design; any provider endpoint can be treated as a selectable substrate in a single system.
If Callosum achieves its technical roadmap, it could become foundational infrastructure for the next generation of AI applications. Success here would accelerate the timeline for downstream companies to build reliable, production-grade AI products. Failure or pivot would signal continued fragmentation in the AI tooling landscape.
“We are scaling heterogeneous compute to unlock a completely new era of AI infrastructure.”
“We co-evolve chips and intelligence together”
“Heterogeneous Intelligence is the defining shift of the next era of AI.”
“Open-source vision-language-action models”
“multi-agent intelligence”
“tool calling”