CostKey

Your agents work all night.Someone should keep watch.

Every agent and workflow in production needs costs, traces and quality checks. CostKey delivers all three out of the box, across 45+ providers, with one line of setup. Track cost, latency, failures and quality scores across every run.

Get started
Every provider. Every agent framework.45+ providers and every major agent framework, in TypeScript or Python. One line to set up, no gateway in between.
Providers
OpenAIAnthropicGeminiAWS BedrockAzure OpenAIMistralDeepSeekxAIGroqTogether AIFireworksOpenRouterCoherePerplexityDeepgramElevenLabsAssemblyAITavilyExaFirecrawlHugging FaceOllama
Agent frameworks
OpenAI Agents SDKLangChainLangGraphVercel AI SDKMastraCrewAILlamaIndexPydantic AIGoogle ADKClaude Agent SDKsmolagentsAutoGenAG2DSStrands AgentsbuOpenTelemetry
Every integration, with setup, in the docs →

Observability

Metrics

Cost, tokens, runs, latency and quality scores across every agent, workflow and model. Compared with the period before.

Agent runs1,20412.4%vs previous1,071
Model cost$1,842.608.1%vs previous$1,704.10
Tokens48.2M5.6%vs previous45.6M
Avg score0.912.2%vs previous0.89
Cost by model
Calls and tokens by model
$1,842.60Total cost
ModelCallsInputOutputCost
claude-sonnet-4-518,42021.4M2.9M$812.40
gpt-59,88012.1M1.6M$541.20
gemini-2.5-flash22,1509.8M0.9M$96.30
gpt-5-mini14,3003.6M0.4M$18.10
deepseek-chat6,2102.2M0.3M$12.40
Usage by agent
Tokens grouped by agent
48.2MTotal tokens
TokensCost
support-agent18.6M
research-agent14.2M
voice-support8.1M
browser-agent5.4M
InputOutput
Scores
Quality checks over time
avg 0.91Across all checks
Over timeSummary
MonTueWedThuFriSatSun
Faithfulness 0.94Tool choice 0.89Answered the question 0.91
Run volume
Runs by agent
1,204Total runs
AgentsWorkflowsTools
support-agent612
research-agent301
voice-support214
browser-agent77
Latency
p50 and p95 per run
4.1 sAvg p50
AgentsWorkflowsTools
MonTueWedThuFriSatSun
p50p95

Traces

Every model call, tool and handoff behind an answer, on one timeline. Six providers in one run, timed and priced together.

voice-supportRun 4b1a · 6 calls · 6 providers
Whole run$0.04326.8 s
0 s2 s4 s6 s
Deepgramtranscribenova-3 · 192 s audio1.1 s$0.0041
OpenAIembedQuerytext-embedding-3-small · 1,240 tokens0.2 s$0.0001
TavilylookupPolicysearch · 1 query0.9 s$0.0080
AnthropicdraftReplyclaude-sonnet-4-5 · 4,180 tokens2.4 s$0.0123
GeminijudgeReplygemini-2.5-flash · 2,050 tokens0.6 s$0.0004
ElevenLabsspeakeleven_v3 · 1,840 characters1.6 s$0.0183
Deepgram 9%OpenAI 0%Tavily 19%Anthropic 28%Gemini 1%ElevenLabs 42%

Down to the line of code.

The function, file and line behind every call. Stop guessing which step spent it.

ContentCost & usageSource
FunctioncrawlAndSummarize
Filesrc/agents/research.ts
Line87
Modelgpt-5
Tokens8,904 in / 1,210 out
Latency6.1 s
Cost$0.0232

Budgets and alerts.

Limits for any agent, workflow, customer or model. Slack hears about it at 80%.

Support monthly$656.00 of $800.0082%
Research daily$9.20 of $20.0046%
Customer acme-corp$2.05 of $5.00 today41%
CostKey APP · 14:02
Allowance warning at 80%
Support monthly: $656.00 of $800.00 this month.
View in CostKey

A kill switch.

When a run hits its cap, the next call is refused before it reaches the provider.

[browser-agent] step 211 · retrying checkout page
[browser-agent] step 212 · retrying checkout page
CostKeyPolicyRefused: budget_exhausted
  Refused before it reached the provider.
CostKey APP · 02:14
A CostKey policy refused a call
Browser run cap refused further matching work (budget_exhausted).
View in CostKey

How do you know the ROI of your agents?
What you spend, and whether you get the outcome.

AgentRunsSpendSucceededPer runPer successWasted
support-agent612$812.4097%$1.33$1.37$23.90
research-agent301$541.2091%$1.80$1.98$48.60
voice-support214$214.8099%$1.00$1.01$2.00
browser-agent77$96.3064%$1.25$1.97$35.00

Per success is spend over the runs that worked, failed runs included. browser-agent looks cheap at $1.25 a run, but only 64% succeed, so each success costs $1.97.

Wasted is what failed runs cost. Quality checks add whether the answers were any good.

If you don’t know this, you are not in control of your budget.

Forecasts. Summaries. Roles.
Everything else you need in production.

Split spend any way.

By customer, feature, function, model, provider or your own label. Price contracts on real numbers.

CustomerFeatureFunctionModelProvider

Which model is letting you down.

Failures, errors, latency and wasted spend for every model you call.

ModelCallsFailedErrorsRetriesp95Wasted
gpt-56,9400.6%429, 500316.4 s$1.88
claude-sonnet-4-518,2040.1%52984.1 s$0.96
gemini-2.5-flash41,3000%—21.2 s$0.04
deepseek-chat7,1000%—03.0 s$0.00
nova-32,3300%—11.1 s$0.00

Month-end, before it arrives.

A forecast from your daily pace, drawn as an estimate. Then ask what 10x the users would cost.

Cumulative spend

ActualEstimated
$0$250$500Day 1Day 10Day 20Day 31$400 budget$248.60 today~$412.90 by day 31
What would more users cost?Today's users$412.902x users$825.8010x users$4,129a month, at today's cost per user

A note on your desk each morning.

Daily and weekly summaries by email or Slack.

Prompts stay private.

Cost, usage, timing and source by default. Prompt storage is a switch.

Store prompts and responses
Cost and usage onlyDefault for new projects
Cost, usage and promptsKept for the days you choose

The whole team, the right view.

Developers see prompts and code. Finance sees costs.

OwnerEverything, with billing
AdminEverything, and the team
DeveloperCosts, prompts and code
ViewerCosts

Self-host

Run it yourself. One command.

Postgres and one container, on your machines. Your data never leaves them.

curl -fsSL https://costkey.dev/install.sh | sh
Self-hosting guide →

Want a hand?

We'll help you set it up and get your first agent watched.

Before you connect.

What is CostKey?

CostKey is open-source observability for AI agents and workflows. It records the cost, latency, failures and quality of every model call and run across 45+ providers, down to the line of code, and adds budgets, Slack alerts and a kill switch. It runs on your own machines.

Which agent frameworks does CostKey work with?

The OpenAI Agents SDK, Vercel AI SDK, Mastra, LangChain and LangGraph, CrewAI, LlamaIndex, Pydantic AI, Google ADK, the Claude Agent SDK, smolagents, AutoGen, AG2, DSPy, Strands and browser-use, each with a one-line integration in TypeScript or Python. Anything that emits OpenTelemetry spans works through the OpenTelemetry bridge, and model calls are captured whatever framework makes them.

Does CostKey sit between my app and my providers?

No. The SDK watches the calls your app already makes and reports them in the background. Requests go straight to your providers, with your keys.

How do I run CostKey?

On your own machines. curl -fsSL https://costkey.dev/install.sh | sh starts Postgres and the CostKey server and console with Docker. Open localhost:4100, sign up, and copy your project's DSN into your app.

How long does setup take?

A few minutes. After the install command, run npx costkey setup for TypeScript or pipx run costkey setup for Python. It signs in to your CostKey, creates a project, adds the DSN to .env and shows the two lines to paste.

Which providers does it work with?

OpenAI, Azure OpenAI, Anthropic, Gemini, Vertex AI, Bedrock, Mistral, DeepSeek, Groq, Together, Fireworks, OpenRouter, Cohere, xAI and more than 25 other model APIs, plus Deepgram, ElevenLabs, AssemblyAI, Cartesia, Tavily, Exa and Firecrawl. You can register any other API in a few lines.

How do I see one agent run as one trace?

Wrap the agent with CostKey.agent and its tools with CostKey.tool (costkey.agent and costkey.tool in Python). Every model call inside lands on that run's timeline with its cost and duration.

How does the kill switch work?

Turn on enforcement in the SDK and set a budget to stop new calls. Before each call, CostKey checks the worst-case cost against the budget and refuses the call before it reaches the provider. Calls already running finish. Supported for 14 major model providers today.

What counts as a failed run?

A run fails when it errors, a model call fails without a successful retry, a tool or step fails, or a judge score you send marks it failing. CostKey shows what each failure cost.

Do you store my prompts?

Not by default. New projects keep cost, usage, timing and source only. Prompt storage is a per-project switch, and you choose how long to keep it. Either way it stays in your own database.

How much does CostKey cost?

Nothing. CostKey is MIT licensed and runs on your machines. Talk to us if you want help setting it up.

Leave a light on.

Get started