TokenHawk
Code-native LLM cost attribution for Rails apps.
Know exactly which feature, domain, or customer is driving your AI bill — without a SaaS dashboard, without changing how you call the API.
30-second demo
# Wrap any LLM call with a domain label
TokenHawk.track(domain: :invoice_extraction) do
client.messages.create(model: "claude-sonnet-4-6", max_tokens: 1024, messages: [...])
end
# That's it. TokenHawk captures tokens, cost, and latency automatically.
Then from the terminal:
$ token_hawk costs --by domain
invoice_extraction $4.21 312 calls
chat_support $1.08 87 calls
pdf_summarizer 38¢ 14 calls
Installation
Add to your Gemfile:
gem "token_hawk"
Run the generator:
bundle exec rails generate token_hawk:install
bundle exec rails db:migrate
Mount the dashboard in config/routes.rb:
mount TokenHawk::Engine, at: "/token_hawk"
Add attribution to your initializer (config/initializers/token_hawk.rb):
TokenHawk.configure do |config|
config.storage = :active_record
end
Tracking calls
With the SDK response (automatic extraction)
TokenHawk.track(domain: :invoice_extraction) do
client.messages.create(...)
end
Supports Anthropic and OpenAI SDK responses out of the box.
With tags (multi-tenant attribution)
TokenHawk.track(domain: :invoice_extraction, tags: { customer_id: 42 }) do
client.messages.create(...)
end
Manual attribution (abstraction layers, PromptCanary, LangChain, etc.)
If your app wraps LLM calls behind a service layer, use record to pass token counts directly:
TokenHawk.record(
domain: :invoice_extraction,
model: "claude-sonnet-4-6",
input_tokens: 820,
output_tokens: 214,
latency_ms: 1340,
tags: { customer_id: 42 }
)
CLI
# Monthly spend by domain (default)
token_hawk costs
# Group by day
token_hawk costs --by day
# Filter to one domain
token_hawk costs --domain invoice_extraction
# Filter to a specific tag value
token_hawk costs --tag customer_id=42
# Group by tag
token_hawk costs --by tag --tag customer_id
# Cost per call by domain (find expensive-per-call domains)
token_hawk efficiency
# Recent calls for debugging
token_hawk recent --limit 20
# JSON output (pipe to jq, etc.)
token_hawk costs --format json | jq .
Dashboard
Mount the engine and visit /token_hawk for a browser-based view:
- Overview — monthly total, top domains by cost, daily spend trend
- Domain detail — per-domain breakdown with top tags and recent calls
- Recent calls — reverse-chronological call log for debugging
- Pricing reference — supported models and their rates
The engine has no auth of its own — protect the mount point with your app's existing authentication:
# config/routes.rb
authenticate :user, ->(u) { u.admin? } do
mount TokenHawk::Engine, at: "/token_hawk"
end
Configuration
TokenHawk.configure do |config|
# Storage backend: :active_record (default) or :memory (tests only)
config.storage = :active_record
# Log telemetry failures instead of raising
config.log_failures = true
config.failure_logger = ->(msg) { Rails.logger.warn(msg) }
# Override or add pricing for unlisted models
config.pricing["my-custom-model"] = { vendor: "anthropic", input: 0.3, output: 1.5 }
end
Subscribe hooks
React to every recorded call in real time:
TokenHawk.subscribe(:call_recorded) do |call|
StatsD.increment("llm.calls", tags: ["domain:#{call.domain}"])
StatsD.gauge("llm.cost_cents", call.total_cost_cents, tags: ["domain:#{call.domain}"])
end
What it doesn't do
- No streaming support — token counts require a complete response
- No real-time dashboard — the UI reads from the database; refresh manually
- No multi-database support — Postgres and SQLite only in v0.1.0
- No alerting — use subscribe hooks to wire your own
Why it exists
Most Rails apps reach for an LLM and start accumulating costs they can't explain. By the time the bill is painful, the usage is spread across dozens of call sites with no attribution. TokenHawk solves this at the source — one wrapper in your code, costs attributed from day one.
License
MIT. See LICENSE.txt.