Your agents work all night.Someone should keep watch.
Every agent and workflow in production needs costs, traces and quality checks. CostKey delivers all three out of the box, across 45+ providers, with one line of setup. Track cost, latency, failures and quality scores across every run.
Observability
Metrics
Cost, tokens, runs, latency and quality scores across every agent, workflow and model. Compared with the period before.
Cost by model
Calls and tokens by model| Model | Calls | Input | Cost |
|---|---|---|---|
| 18,420 | 21.4M | $812.40 | |
| 9,880 | 12.1M | $541.20 | |
| 22,150 | 9.8M | $96.30 | |
| 14,300 | 3.6M | $18.10 | |
| 6,210 | 2.2M | $12.40 |
Usage by agent
Tokens grouped by agentScores
Quality checks over timeRun volume
Runs by agentLatency
p50 and p95 per runTraces
Every model call, tool and handoff behind an answer, on one timeline. Six providers in one run, timed and priced together.
Down to the line of code.
The function, file and line behind every call. Stop guessing which step spent it.
crawlAndSummarizesrc/agents/research.tsBudgets and alerts.
Limits for any agent, workflow, customer or model. Slack hears about it at 80%.
A kill switch.
When a run hits its cap, the next call is refused before it reaches the provider.
How do you know the ROI of your agents?
What you spend, and whether you get the outcome.
| Agent | Spend | Succeeded | Per success |
|---|---|---|---|
| support-agent | $812.40 | 97% | $1.37 |
| research-agent | $541.20 | 91% | $1.98 |
| voice-support | $214.80 | 99% | $1.01 |
| browser-agent | $96.30 | 64% | $1.97 |
Per success is spend over the runs that worked, failed runs included. browser-agent looks cheap at $1.25 a run, but only 64% succeed, so each success costs $1.97.
Wasted is what failed runs cost. Quality checks add whether the answers were any good.
If you don’t know this, you are not in control of your budget.
Forecasts. Summaries. Roles.
Everything else you need in production.
Split spend any way.
By customer, feature, function, model, provider or your own label. Price contracts on real numbers.
Which model is letting you down.
Failures, errors, latency and wasted spend for every model you call.
| Model | Failed | p95 | Wasted |
|---|---|---|---|
| 0.6% | 6.4 s | $1.88 | |
| 0.1% | 4.1 s | $0.96 | |
| 0% | 1.2 s | $0.04 | |
| 0% | 3.0 s | $0.00 | |
| 0% | 1.1 s | $0.00 |
Month-end, before it arrives.
A forecast from your daily pace, drawn as an estimate. Then ask what 10x the users would cost.
Cumulative spend
ActualEstimatedA note on your desk each morning.
Daily and weekly summaries by email or Slack.
Daily spend · 19 Oct
$281.40 yesterday, 12% more than the day before.
- Top features support-agent, research-agent, summarize
- Top models claude-sonnet-4-5, gpt-5, gemini-2.5-flash
- Budgets Support monthly is at 82%
Prompts stay private.
Cost, usage, timing and source by default. Prompt storage is a switch.
The whole team, the right view.
Developers see prompts and code. Finance sees costs.
| Owner | Everything, with billing |
| Admin | Everything, and the team |
| Developer | Costs, prompts and code |
| Viewer | Costs |
Self-host
Run it yourself. One command.
Postgres and one container, on your machines. Your data never leaves them.
curl -fsSL https://costkey.dev/install.sh | shBefore you connect.
What is CostKey?
CostKey is open-source observability for AI agents and workflows. It records the cost, latency, failures and quality of every model call and run across 45+ providers, down to the line of code, and adds budgets, Slack alerts and a kill switch. It runs on your own machines.
Which agent frameworks does CostKey work with?
The OpenAI Agents SDK, Vercel AI SDK, Mastra, LangChain and LangGraph, CrewAI, LlamaIndex, Pydantic AI, Google ADK, the Claude Agent SDK, smolagents, AutoGen, AG2, DSPy, Strands and browser-use, each with a one-line integration in TypeScript or Python. Anything that emits OpenTelemetry spans works through the OpenTelemetry bridge, and model calls are captured whatever framework makes them.
Does CostKey sit between my app and my providers?
No. The SDK watches the calls your app already makes and reports them in the background. Requests go straight to your providers, with your keys.
How do I run CostKey?
On your own machines. curl -fsSL https://costkey.dev/install.sh | sh starts Postgres and the CostKey server and console with Docker. Open localhost:4100, sign up, and copy your project's DSN into your app.
How long does setup take?
A few minutes. After the install command, run npx costkey setup for TypeScript or pipx run costkey setup for Python. It signs in to your CostKey, creates a project, adds the DSN to .env and shows the two lines to paste.
Which providers does it work with?
OpenAI, Azure OpenAI, Anthropic, Gemini, Vertex AI, Bedrock, Mistral, DeepSeek, Groq, Together, Fireworks, OpenRouter, Cohere, xAI and more than 25 other model APIs, plus Deepgram, ElevenLabs, AssemblyAI, Cartesia, Tavily, Exa and Firecrawl. You can register any other API in a few lines.
How do I see one agent run as one trace?
Wrap the agent with CostKey.agent and its tools with CostKey.tool (costkey.agent and costkey.tool in Python). Every model call inside lands on that run's timeline with its cost and duration.
How does the kill switch work?
Turn on enforcement in the SDK and set a budget to stop new calls. Before each call, CostKey checks the worst-case cost against the budget and refuses the call before it reaches the provider. Calls already running finish. Supported for 14 major model providers today.
What counts as a failed run?
A run fails when it errors, a model call fails without a successful retry, a tool or step fails, or a judge score you send marks it failing. CostKey shows what each failure cost.
Do you store my prompts?
Not by default. New projects keep cost, usage, timing and source only. Prompt storage is a per-project switch, and you choose how long to keep it. Either way it stays in your own database.
How much does CostKey cost?
Nothing. CostKey is MIT licensed and runs on your machines. Talk to us if you want help setting it up.