Skip to content

Clawsseum is in early access. Request access

Platform

One platform from first prompt to production.

The SDK, adapters, evals and runtime share one definition of your agent, so the code you run locally is the code you test and the code you ship. Use all four or start with one.

Architecture

How a request moves through Clawsseum

Channels send events, the runtime runs your agent, adapters handle every call to an outside system, and evals use the recorded traffic to test the next version.

Channels

S

Slack

threads, approvals

WW

Web widget

your product

Gmail

inbound mail

LiveKit

voice

Runtime

healthy

Agent

support-agent@v13

planmemory: casetools: 4budget $0.05
approvalsredactiontraces

Adapters

Models

O

OpenAI

primary

Google Gemini

fallback

Apps

Zendesk

tickets:read

Stripe

refunds:write

Data

PostgreSQL

read replica

Evals

record trafficreplay v13 vs v12scoredeploy or block

sandboxed side effects

01Agent SDK

Agents as code.

Agents are ordinary TypeScript or Python modules. Review them in a pull request, test them in CI and move them between environments. No visual builder to export, no hidden state.

agents/support_agent.py
from clawsseum import agent, tool
from clawsseum.adapters import openai, stripe
from clawsseum.adapters import slack, zendesk
from pydantic import BaseModel

class Refund(BaseModel):
    charge_id: str
    amount: float

@tool(side_effect=True)  # requires approval at runtime
async def issue_refund(args: Refund, ctx):
    return await ctx.adapters.stripe.refunds.create(
        charge=args.charge_id, amount=args.amount
    )

support_agent = agent(
    name="support-agent",
    model=openai("gpt-5"),
    adapters=[
        zendesk(),
        stripe(scope=["refunds:write"]),
        slack(),
    ],
    tools=[issue_refund],
    memory={"kind": "case", "ttl": "30d"},
)

Typed tools

Inputs and outputs are validated with zod or pydantic before anything reaches an adapter.

Durable runs

Runs survive deploys and crashes, and can wait days for a human or a webhook.

Scoped memory

Case, user and org memory with TTLs, so context never leaks between customers.

Readable plans

Plans are plain data. Assert on them in tests and diff them between versions.

Local runtime

Hot reload against live adapters or recorded fixtures on your laptop.

Sub-agents

Hand work to specialist agents with their own budgets, tools and scopes.

02Adapters

One interface for every system.

Each of the 102 adapters, whether it wraps a model, a channel, an app or a database, uses the same typed interface: actions with schemas, declared side effects, scopes, idempotency and redaction. Agents call actions, never raw APIs.

stripe.refunds.create

action
input
{ charge: string, amount: number, reason?: string }
output
Refund { id, status, amount }
scopes
refunds:write
side effect
yes, approval above policy limit
idempotency
key = run.id + charge
retries
3, exponential, 429 and 5xx only
rate limit
provider default, shared per account
redact
customer.email, card.last4
recorded
yes, replayable in evals
Same shape for Zendesk, Salesforce, Postgres, Slack and every other adapter.

Scoped credentials

Secrets live in a vault and are issued per run with least-privilege scopes.

Model routing

Set primary and fallback models per agent and change them in config or from eval results.

Retries and rate limits

Retries, rate limits, timeouts and circuit breakers tuned for each provider.

Redaction

Mask sensitive fields before a model or a trace ever sees them.

Recordings

Every call is recorded, so evals can replay it without touching production.

Bring your own

Wrap MCP servers, OpenAPI specs, GraphQL and gRPC services in minutes.

03Evals

Catch regressions before your customers do.

Datasets come from real production traffic, not hand-written happy paths. Each new version runs every case side by side with the current version and graders score both. The result is a number your CI can enforce.

evals.config.ts
import { evals, grader, check } from "@clawsseum/sdk";

export default evals({
  dataset: "support/refunds", // 412 production cases
  current: "support-agent@v12",
  candidates: ["support-agent@v13"],
  graders: [
    grader.rubric("./rubrics/resolution.md", {
      model: "gemini-2.5-pro",
    }),
    check.noPII(),
    check.policy("refunds", { maxWithoutApproval: 500 }),
    check.tone({ forbid: ["blame", "legalese"] }),
  ],
  releaseCheck: {
    minWinRate: 0.55,
    maxCriticalRegressions: 0,
    maxCostDelta: "+10%",
  },
});

Datasets from production

Turn real conversations into test cases, with PII redacted before storage.

Synthetic variations

Expand one test case into paraphrases, other languages and edge cases.

Graders and checks

Model graders for quality, deterministic checks for policy, pairwise comparison for preference.

Human review

Close calls and grader disagreements go to a reviewer instead of being guessed.

Quality score

One score per dataset, so versions, prompts and models can be compared directly.

Checks in CI

A required check on every pull request. Versions that score worse do not merge.

04Runtime

Run agents safely in production.

The runtime sits between your agent and the systems it touches. It controls what an agent may do, what needs a human, how much it can spend and how a new version rolls out.

runtime.policy.ts
import { runtime } from "@clawsseum/sdk";

export default runtime.policy({
  agent: "support-agent",
  approvals: [
    {
      tool: "issue_refund",
      when: "amount > 500",
      route: "slack:#support-leads",
    },
    { adapter: "gmail", action: "send", route: "owner" },
  ],
  budgets: { perRun: "$0.05", perDay: "$40" },
  redact: ["card.number", "customer.address"],
  rollout: {
    canary: [10, 50, 100],
    rollbackOn: { errorRate: 0.02 },
  },
});

Approvals

Send risky actions to Slack or the dashboard with full context and a one-click decision.

Budgets and kill switches

Cap spend per run, per agent and per day. Stop an agent everywhere at once.

PII redaction

The same policies apply to every adapter call and every trace.

Tracing

Every step, prompt, tool call and cost is searchable and exports over OpenTelemetry.

Gradual rollouts

Canary new versions and roll back to the current version on errors or drift.

Your cloud

Run the runtime in your own VPC on Enterprise when data cannot leave it.

Open standards

No lock-in.

Agents are code in your repo, datasets export as JSON, traces go out over OpenTelemetry, and adapters use open protocols. If you leave, you take everything with you.

Model Context Protocol

Any MCP server is an adapter.

OpenAPI

Specs become typed adapters.

A

Agent2Agent

Delegate to agents on other platforms.

O

OpenTelemetry

Send traces to your existing stack.

TypeScript

First-class SDK and CLI.

Python

Same interface, same evals.

Already have agents?

Run evals on them today.

Wrap an existing agent from any framework with one function call and start scoring releases this week. Move to the SDK and runtime when you are ready, or not at all.