Company
We are building the platform for reliable AI agents.
Building an agent is now the easy part. Knowing it works, and keeping it working as models, tools and customers change, is not. That is the problem we work on.
Why we exist
Most teams we talk to have a demo that works and a production agent they do not fully trust. The gap is not model quality. It is testing, integrations and control.
Models change monthly.
A better model ships and a prompt that worked last week quietly breaks. Teams pin old versions out of fear and leave quality on the table. Swapping a model should be a change you can measure, not a leap of faith.
Integrations are glue.
The hard part of an agent is rarely the reasoning. It is auth, retries, pagination, rate limits and the twelfth slightly different CRM. Most agent code is glue, and glue is where incidents live.
Regressions are silent.
Agents do not crash when they get worse. They get a little ruder, a little slower, a little more generous with refunds. You find out from a customer screenshot, not a test.
Clawsseum closes that gap with four parts that work together: the Agent SDK to build, Adapters to connect to your stack, Evals to test every release and the Runtime to run agents safely in production.
Principles
Our principles.
Six rules we use to settle arguments, in the product and in the company.
- 01
Measure every change.
If a change cannot be measured, it does not ship. That includes ours: every Clawsseum release runs through our own evals.
- 02
No lock-in.
Every model and system sits behind the same interface. Leaving us should be as easy as switching providers.
- 03
Humans approve risky actions.
Agents propose; people approve anything that moves money or data. Autonomy is granted per tool, not per agent.
- 04
No markup on tokens.
You pay model providers directly at their prices. We charge for the platform, never a markup on tokens.
- 05
Reliable infrastructure first.
Durable execution, retries and idempotency keys come before features. The platform should be predictable so your agents can be.
- 06
Small surface, clear docs.
Fewer concepts, each explained completely. If you need a forum thread to use a feature, the feature needs work.
How we work
Small team. Short loops.
- Written by default
- Remote-first. Decisions land in documents, not in meetings, so anyone can catch up and push back.
- Weekly releases
- We ship to early-access teams every week, and every release passes our own evals first.
- Partners shape the roadmap
- Early-access teams get a direct line to the engineers building what they use.
- Honest about the edges
- We say what is GA, what is in beta and what is not built yet. The roadmap below is the same list we use.
Roadmap
What is shipped and what is next.
The same list we work from. Dates move; order rarely does. Need something that is not here? Ask for it.
Now
In early accessAgent SDK
TypeScript and Python, with typed tools, memory and the local runtime.
Core adapters
Every model adapter plus the core channel, app and data adapters.
Evals and CI checks
Build datasets from production, score new versions and block releases in CI.
Hosted runtime
Durable execution, approvals, budgets and traces.
Next
In progressA2A delegation, GA
Hand work to agents built on other platforms, with the same scopes.
Voice agents, GA
Realtime speech with barge-in, on LiveKit, Deepgram and ElevenLabs.
EU region
Traces, memory and eval datasets processed and stored in the EU.
Microsoft Teams, GA
Approvals and conversations inside Teams, with adaptive cards.
Later
PlannedSelf-hosted control plane, GA
The whole platform inside your own cloud, not only the runtime.
Adapter marketplace
Publish and install adapters built by partners and the community.
On-device runtime
Run small agents at the edge and on laptops, with the same SDK.
Missing an adapter your agents need? We build the most-requested ones first.
Request an adapterCareers
We are hiring.
We are not posting a wall of roles. If you recognize yourself below, write to us with something you built and why you are proud of it. A link beats a resume.
Product engineers
TypeScript, distributed systems and the taste to make hard things feel small.
Applied AI researchers
Evaluation, model graders and agent reliability, measured on real traffic.
A founding designer
Interfaces for software that thinks, from the CLI to the trace viewer.