Step 1
Install
You need Node 20 or later for TypeScript, or Python 3.11 or later. Everything ships in one package: @clawsseum/sdk for TypeScript and clawsseum for Python, with agents, tools, adapters and evals.
# TypeScript (Node 20+)
npm install @clawsseum/sdk
# or
pnpm add @clawsseum/sdk
# Python (3.11+)
pip install clawsseum
# The CLI
npm install -g clawsseum
clawsseum loginScaffold a project from the support template. It gives you a working agent, a prompt and a starter eval dataset.
$ clawsseum init support-agent --template support
Created clawsseum.config.ts
Created agents/support-agent.ts
Created prompts/support-agent.md
Created evals.config.ts
Created evals/datasets/support-refunds.jsonl (24 starter cases)
Next: clawsseum adapters add zendesk stripe slackProject settings live in clawsseum.config.ts: which agents to load, your environments, where secrets are stored and where traces go.
import { defineConfig } from "@clawsseum/sdk";
export default defineConfig({
project: "acme-support",
agents: ["./agents/*.ts"],
environments: {
staging: { region: "us-east" },
production: { region: "us-east", releaseCheck: "required" },
},
// "clawsseum" | "aws-secrets-manager" | "gcp-secret-manager" | "vault"
secrets: { provider: "clawsseum" },
telemetry: { otlpEndpoint: process.env.OTEL_EXPORTER_OTLP_ENDPOINT },
});Step 2
Create an agent
An agent is a model, a set of adapters, the tools built on them and a prompt. Tools are typed: the schema you write is the schema the model sees, and inputs are validated before your code runs.
import { agent, tool, z } from "@clawsseum/sdk";
import { google, openai, slack, stripe, zendesk } from "@clawsseum/sdk/adapters";
const issueRefund = tool({
name: "issue_refund",
description: "Refund a charge, fully or partially.",
input: z.object({
chargeId: z.string(),
amountCents: z.number().int().positive(),
reason: z.enum(["duplicate", "fraudulent", "requested_by_customer"]),
}),
sideEffect: true, // subject to runtime approval policies
run: ({ chargeId, amountCents, reason }, { adapters }) =>
adapters.stripe.refunds.create({ charge: chargeId, amount: amountCents, reason }),
});
export default agent({
name: "support-agent",
model: openai("gpt-5"),
fallback: [google("gemini-2.5-pro")],
adapters: [
zendesk({ scope: ["tickets:read", "tickets:write"] }),
stripe({ scope: ["charges:read", "refunds:write"] }),
slack({ channels: ["#support-escalations"] }),
],
tools: [issueRefund],
instructions: "./prompts/support-agent.md",
memory: { kind: "case", ttl: "30d" },
limits: { maxSteps: 40, maxCostUsd: 0.25 },
});from typing import Literal
from pydantic import BaseModel, PositiveInt
from clawsseum import agent, tool
from clawsseum.adapters import google, openai, slack, stripe, zendesk
class RefundInput(BaseModel):
charge_id: str
amount_cents: PositiveInt
reason: Literal["duplicate", "fraudulent", "requested_by_customer"]
@tool(side_effect=True)
async def issue_refund(input: RefundInput, ctx) -> dict:
"""Refund a charge, fully or partially."""
return await ctx.adapters.stripe.refunds.create(
charge=input.charge_id, amount=input.amount_cents, reason=input.reason
)
support_agent = agent(
name="support-agent",
model=openai("gpt-5"),
fallback=[google("gemini-2.5-pro")],
adapters=[
zendesk(scope=["tickets:read", "tickets:write"]),
stripe(scope=["charges:read", "refunds:write"]),
slack(channels=["#support-escalations"]),
],
tools=[issue_refund],
instructions="./prompts/support-agent.md",
)Run it locally. The playground shows every step: model calls, tool calls, adapter requests and their latency.
# Local playground with hot reload and live traces
clawsseum dev support-agent
# One-off run that prints the full trace
clawsseum run support-agent --input "Customer 4821 was charged twice for March"Concept
Adapters
Every adapter, built-in or your own, implements the same interface. That is what lets you swap Slack for Teams or one model for another without touching agent code.
- Typed actions. Each adapter exposes actions with input and output schemas. The agent only sees actions inside the scopes you grant.
- Scoped credentials. Secrets stay in your secret store and are issued per run. Scopes are enforced on every call, not in the prompt.
- Retries and rate limits. Idempotency keys, backoff and provider limits are handled for you. Retries are not billed as steps.
- Redaction. Fields listed in
redactare masked before the model sees them and in every trace. - Recording. Every call is recorded with its response and latency, so evals can replay real traffic.
$ clawsseum adapters add zendesk stripe slack
zendesk tickets:read tickets:write connected (OAuth)
stripe charges:read refunds:write connected (restricted key)
slack chat:write #support-escalations connected (OAuth)
Credentials stored in Clawsseum secrets. Scopes enforced on every call.For systems outside the catalog, wrap an MCP server with adapter.fromMCP, an OpenAPI spec with adapter.fromOpenAPI, or define actions by hand:
import { adapter, z } from "@clawsseum/sdk";
export const inventory = adapter.define({
name: "inventory",
auth: adapter.auth.bearer("INVENTORY_API_KEY"),
actions: {
getStock: {
input: z.object({ sku: z.string() }),
output: z.object({ sku: z.string(), available: z.number().int() }),
run: async ({ sku }, { http }) => http.get("/stock/" + sku).json(),
},
},
retry: { attempts: 3, backoff: "exponential" },
rateLimit: { perSecond: 10 },
redact: [],
});Models are adapters too. Overrides let you try a different model per environment while the agent definition stays put.
// clawsseum.config.ts: try a different model in staging, agent code unchanged
import { mistral } from "@clawsseum/sdk/adapters";
export default defineConfig({
// ...
overrides: {
staging: {
"support-agent": { model: mistral("mistral-large") },
},
},
});Concept
Evals: datasets, graders and release checks
Evals answer one question before every release: is the new version better than the one in production? They run a dataset of test cases against the new version and the current version, have graders score both, and apply your release check thresholds.
The best test cases come from real traffic. Record a sample from production; it is redacted with your adapter rules.
# Turn 5% of production conversations from the last 14 days into test cases.
# PII is redacted with the same rules your adapters use.
clawsseum eval record --agent support-agent --env production \
--sample 5% --days 14 --dataset support/refundsGraders combine rubric scoring by a model with deterministic checks: no PII in replies, policy compliance, step limits. Replays use recorded adapter responses, and side effects go to sandboxes.
import { evals, check, grader } from "@clawsseum/sdk";
import { google } from "@clawsseum/sdk/adapters";
export default evals({
dataset: "support/refunds", // 412 cases
current: "support-agent@v12",
candidates: ["support-agent@v13"],
replay: { adapters: "recorded", sideEffects: "sandbox" },
graders: [
grader.rubric("./rubrics/resolution.md", { model: google("gemini-2.5-pro"), weight: 0.6 }),
check.noPII(),
check.policy("./policies/refunds.yaml"),
check.maxSteps(30),
],
releaseCheck: {
minWinRate: 0.55,
maxCriticalRegressions: 0,
maxCostDelta: "+10%",
maxLatencyP95Delta: "+15%",
},
});$ clawsseum eval run --dataset support/refunds --against v12
Eval support/refunds 412 cases current v12 new v13
results 412/412 better 263 tie 91 worse 58
win rate 63.8%
regressions 0 critical, 3 minor (tone)
cost / task $0.018 -> $0.014 (-22%)
p95 latency 4.1s -> 3.2s
Release check PASS v13 can be deployedPut the same command in CI. With --check the job fails when a threshold is missed, and the pull request shows the eval summary.
# .github/workflows/evals.yml
name: evals
on:
pull_request:
paths: ["agents/**", "prompts/**", "evals.config.ts"]
jobs:
evals:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- run: npm ci
- run: npx clawsseum eval run --dataset support/refunds --against production --check
env:
CLAWSSEUM_TOKEN: ${{ secrets.CLAWSSEUM_TOKEN }}Step 3
Deploy: approvals, budgets and rollouts
The runtime runs your agents in production. Runs are durable: an agent waiting on an approval or a slow API holds no compute and survives deploys. Policies decide which tool calls need a human, and budgets cap what any run or day can spend.
// clawsseum.config.ts
export default defineConfig({
// ...
runtime: {
approvals: [
{
tool: "issue_refund",
when: "input.amountCents > 50000",
route: "slack:#support-escalations",
timeout: "4h",
},
],
budgets: {
perRun: { usd: 0.25, steps: 40 },
perDay: { usd: 300 },
},
killSwitch: { errorRate: 0.05, window: "5m" },
},
});Deploying runs the release check, builds the version, issues scoped secrets and rolls out behind a canary.
$ clawsseum deploy support-agent --env production
Release check support/refunds v13 vs v12 PASS 63.8% win, 0 critical
Build support-agent@v13 1 tool, 3 adapters
Secrets zendesk, stripe, slack scoped, issued per run
Rollout canary 10% -> 50% -> 100% auto-rollback if errors > 2%
Live support-agent@v13 in productionEvery run produces a trace with model calls, tool calls, approvals and cost. Traces export over OpenTelemetry to the observability stack you already run.
Reference
CLI reference
| Command | What it does |
|---|---|
| clawsseum login | Authenticate the CLI with your workspace. |
| clawsseum init <name> [--template] | Scaffold an agent, its prompt, a config and a starter eval dataset. |
| clawsseum dev <agent> | Local playground with hot reload and live traces. |
| clawsseum run <agent> --input <text> | Run once and print the full trace. |
| clawsseum adapters add <names...> | Install adapters and request scoped credentials. |
| clawsseum adapters ls | List adapters with their scopes and health. |
| clawsseum eval record | Turn production traffic into redacted test cases. |
| clawsseum eval run [--check] | Run a dataset against the current version and enforce thresholds. |
| clawsseum eval scores | Show the quality score of every version. |
| clawsseum deploy <agent> --env <env> | Release check, build, scoped secrets and a canary rollout. |
| clawsseum rollback <agent> --to <version> | Return to a previous version immediately. |
| clawsseum logs <agent> [--follow] | Stream runs, tool calls and approvals. |
Every command accepts --json for scripting and --env to target an environment. Run clawsseum help <command> for flags.
Next