SpendFriend

AI Spend Control Plane

Meter every token. Cap every dollar. Block no one.

One gateway across Claude, ChatGPT, Gemini, and Bedrock. Per-employee budgets that degrade gracefully to cheaper models — your team never hits a wall.

See it in action
Meters acrossAnthropicOpenAIGeminiAmazon Bedrock

AI spend · October

$48,211

+12.4% vs Sept
OpenAI$21,412 · 44%
Anthropic$14,806 · 31%
Cursor$7,590 · 16%
Bedrock$4,403 · 9%

Quota event: Engineering hit 90% of premium allowance — auto-routed to efficient models. Zero downtime.

The problem

Four gaps no provider dashboard will close

Organizations are spending on AI faster than they can track it. These are the gaps SpendFriend was built to close.

18%

Nobody can attribute it

AI spend arrives on three rails — SaaS seats, API tokens, and cloud-hosted infra. Roughly 18% of it can't be tied to any team or outcome.

11%

Nobody can forecast it

Only ~11% of businesses can accurately forecast AI spend. Classic cloud FinOps tools weren't built for LLM economics.

32%

Cheaper isn't cheaper

In ~32% of enterprise use cases, the cheaper model costs more — once retries, longer chains, and failed first-pass answers are counted.

78%

It's hiding in other tools

Usage-priced AI add-ons ship inside SaaS IT already pays for — up to 78% of IT leaders report surprise charges from these hidden features.

How it works

From scattered AI spend to one controlled ledger

01

Issue virtual keys

Every employee gets a per-person key that maps to provider credentials held in your vault. They paste it into any OpenAI-compatible tool — or use the built-in chat.

02

Meter every request

Tokens, cost, latency, and cache hits land in one ledger — whether the request came through our dashboard or a passthrough API key in someone's own tool.

03

Enforce budgets, never block

Set allowances at org, team, or employee level. When premium quota runs out, requests degrade to mid-tier or free models instead of stopping.

04

Optimize on real numbers

The routing advisor replays your prompt history and prices cost per successful outcome per model — so every routing change comes with a dollar figure attached.

Metering

Every request metered — whatever surface it came from

Requests flow through the SpendFriend gateway — via the built-in dashboard chat, or per-employee passthrough keys your team pastes into the tools they already use. Token counts, cost, latency, and cache hits append to one ledger.

Metering is metadata-first: prompt and response bodies are stored only if your org opts in.

Daily consumption · all providers

Refreshed daily
Oct 1Oct 14

Budgets & Fallback

Cap the spend. Never block the work.

Set premium allowances at org, team, or employee level — daily, weekly, or monthly. When quota runs out, requests don’t stop: they degrade transparently to mid-tier or free models, mid-conversation.

Employees see a badge — “Running on efficient mode” — and keep going. Admins keep hard-block available as the exception, not the rule.

Engineering · monthly premium allowance

Efficient mode

$4,982 / $5,000

Premium · Claude Opus, GPT-4-classExhausted
Mid · Sonnet / GPT-4o-mini-classActive now
Free · Flash / self-hostedStandby

Quota crossed → transparent downgrade mid-conversation. The employee keeps working; the budget holds.

Routing Advisor

The cheapest model is the one that finishes the job

Cost per request lies. The advisor replays your actual prompt history and measures cost per successful outcome per model — retries, long chains, and failed first passes included.

Recommendations arrive with real numbers — “route this workload, save $Y/mo” — and nothing changes until an admin approves it.

Cost per successful outcome · support drafts

claude-sonnet-4.5$0.31 Recommended
gpt-5-mini$0.44 Retries incl.
claude-opus-4.1$0.58

Recommendation: route “support drafts” to Sonnet — est. save $4,120/mo.

Awaits admin approval

Consolidation

Three rails of AI spend. One ledger.

Gateway-metered API tokens, SaaS seat billing, and cloud spend from Bedrock and Vertex — normalized to the same org, team, and workflow so anyone who needs the number sees the same number.

Finance, IT, and engineering stop reconciling partial spreadsheets that each claim to be the AI budget.

Total AI spend · one ledger

$48,211

API tokens · gateway-metered ✓SaaS seats · billing ingestion ✓Cloud billing · AWS / GCP ✓Shadow AI · flagged events ✓+ whatever you add next

Every spend source normalized to org → team → member → workflow, so the ledger answers “what does the support team cost” no matter which rail the spend arrived on.

Leak Audit

Find the spend hiding in context windows

Provider logs show what was spent, not why. SpendFriend tracks cache-hit rates and context size per request, then flags the agents stuffing un-cached blocks into context windows.

You get a per-application leak report quantifying wasted spend — not just a spike on a chart.

Context & cache leak report · this week

2 leaks found
Cache-hit rate · doc-qa agent41%
Context growth · research agent+214k tok / session
Requests over 100k context312 this week

doc-qa re-sends a 90k-token spec uncached on every call — $1,870/mo wasted.

Shadow AI

Catch the AI features hiding inside tools you already pay for

Salesforce, Zoom, Notion, and Jira all ship usage-priced AI add-ons. A lightweight browser extension flags billable AI consumption outside sanctioned channels and feeds it into the same ledger.

It observes request patterns — never prompt content — and hands IT a review queue, not a surveillance feed.

Shadow AI candidates · pending IT review

Salesforce Einstein prompts41 billable eventsunattributed
Notion AI queries18 billable eventsunattributed
Zoom AI Companion summaries27 billable eventsunattributed

Observed by the browser extension — usage patterns only, no prompt content captured.

Unit Economics

Turn tokens into a P&L line

Meter events carry workflow tags — ticket resolved, contract reviewed, asset generated — so reporting computes cost per outcome, benchmarked against the human labor it replaced.

“AI spend went up” becomes “support cost per ticket dropped 38%.”

Cost per outcome · workflow tags

Support ticket resolved$0.42 vs $6.10 human baseline
Contract reviewed$1.18 vs $220 human baseline
Marketing asset generated$0.87 vs $95 human baseline

Tags attach to meter events — token totals become a P&L line finance can actually read.

Governance

Enterprise controls, built in — not bolted on

SSO / SAML + SCIM

Provision employees into your org → team hierarchy from your IdP. RBAC for admins, team managers, finance viewers, and employees.

Virtual keys & revocation

Employees never touch provider credentials. Per-person keys make offboarding and incident response instant — one revoke, one person.

Policy & audit trail

Budgets, request-class rules, and routing policies are versioned and audit-logged, because procurement and compliance will ask.

Privacy-first metering

Token counts and cost are logged by default; prompt bodies only if your org opts in. Custom endpoints and DLP hooks on Enterprise.

One control plane for finance, IT & team leads

For Finance

A continuously refreshed ledger across every provider, cost per workflow outcome, and forecasts built on actuals — not invoice archaeology.

See finance solutions

For IT & Platform

One gateway, per-employee virtual keys, budgets at every scope, and instant revocation — with an audit trail on every policy change.

See IT solutions

For Team Leads

Set allowances your team can't blow through and they'll never feel — premium quota exhausts into cheaper models, not dead ends.

See how fallback works

See every AI dollar. Control where it goes.

Connect the gateway in minutes. Per-seat platform pricing plus metered consumption — and employees who never hit a wall.

See it in action