AI FinOps & Agent Guardrails

Stop runaway agent loops and surprise AI bills before they reach your invoice.

Cognocient sits in front of your LLM calls to enforce pre-call budget caps, stop agents stuck repeating the same tool call, and attribute every dollar by feature, team, and session. No code rewrite: change one base URL.

Start your 10-day free trial

Every feature · No credit card · Then free forever on 1 provider

Import a usage export for a waste check

CSV or OpenTelemetry export · No code change or production access

Works with any OpenAI-compatible SDK or framework, including for Claude, Gemini, Bedrock and Vertex models. Keep your tracing: Cognocient can export every call to Langfuse or any OpenTelemetry backend. Read the 2-minute quickstart →

One request, watched end to end
Your appPOST /v1/chat/completions

Cognocient proxy

Budget check passed
Attribution tagged
Running cost$0.0000
AI providerresponse returned

Attributed automatically

Feature
checkout-agent
Team
Engineering
Session
sess_8f2e…
Engineers

Stop looping agents and leaked keys in the request path, and see which feature and model every dollar went to.

Finance teams

Budget vs. actuals, chargeback by GL account and board-ready PDF reports in one click, without a spreadsheet.

Founders & CFOs

Hard budget caps that hold at the proxy, so one bad night can't take out the quarter's AI budget.

Why AI bills get out of hand

Three ways an AI bill blows up overnight.

Each one is invisible until the invoice arrives. Here is what Cognocient puts in front of each.

01

An agent gets stuck in a loop.

It calls the same tool with the same arguments, or keeps hitting errors, all night. Nothing crashes, so nothing alerts.

What stops it

  • Failure Loop Breaker stops the run· Growth+
  • Per-run dollar caps with X-Cost-Run-ID
  • max_tokens clamped to the budget left
02

A key leaks, or a deploy goes wrong.

Spend jumps from dollars an hour to dollars a minute, and the first sign is the provider's bill.

What stops it

  • Dollar-per-minute velocity limits
  • Alerts on a new source IP or model family per key· Base+
  • One-click Emergency Freeze of all proxied spend· Every plan
03

Nobody can say what is driving the bill.

The invoice shows one number per provider. Finance asks which feature or team it belongs to, and engineering has to guess.

What stops it

  • Spend by feature, team and user from request headers
  • Budgets per feature, department, user or key that alert, degrade or block
  • Board-ready PDF reports· Base+

In the request path

Autonomous agents shouldn't have unlimited access to your credit card.

Tracing tools show you an agent looped after the bill has landed. Cognocient sits in the request path, so it checks each call before it reaches the provider and can stop it there.

Failure Loop Breaker

The problem

An agent keeps calling the same tool with the same arguments, or keeps hitting errors, turn after turn, and never finishes.

What Cognocient does

Within a session, trips on 3 identical tool calls in a row, or a run of consecutive errors: 2 that retrying won't fix (bad input, auth), 3 it can't classify, or 5 retryable ones (timeouts, 429s, 5xx). Checked before the next call is sent. Choose alert-only or stop the run.

Growth and up · needs the X-Cost-Session header

Pre-call max_tokens clamp

The problem

A max_tokens limit in application code only protects the code paths that remember to set it.

What Cognocient does

When a call reserves against a hard-cap budget, Cognocient lowers the request's max_tokens so the response can't cost more than the budget has left, and tells you in a response header.

Applies to Block-mode budgets

Reads degrade, writes stop

The problem

A hard stop in the middle of an agent run can strand a user request that was nearly done.

What Cognocient does

When a budget runs out, calls whose tools look like writes (create, delete, send, deploy…) are hard-stopped. Read-only tool calls continue on a cheaper model from the same family where one is mapped, e.g. Claude Sonnet to Haiku.

Detected from tool names, or set explicitly with X-Cost-Workload

How it works

From black box to boardroom clarity

One base URL change. Every API call observed, every dollar explained.

2 min setup
Step 01

Point your API calls at Cognocient

Change one base URL. Your AI traffic flows through the proxy: OpenAI, Anthropic, Gemini, Bedrock, Vertex and 5 more provider types.

Quickstart
Optional headers
Step 02

Every call is tagged and attributed

Add X-Cost-Feature and X-Cost-Session headers. Spend is sliced by feature, team, session, and user automatically.

All attribution headers
Automatic
Step 03

Waste and spikes get flagged.

Cost anomalies are flagged on every plan; retries, over-sized models and growing context on Base and up, with one-click recommendations on Growth and up.

What gets flagged
app.pyThe whole integration
from openai import OpenAI

client = OpenAI(
    api_key="sk-cog-...",                       # your Cognocient proxy key
    base_url="https://api.cognocient.com/v1",   # the one-line change
    default_headers={"X-Cost-Feature": "support-agent"},  # optional attribution
)

# Same SDK, any provider: OpenAI, Claude, Gemini, Bedrock, Vertex...
client.chat.completions.create(
    model="claude-sonnet-4-6",
    messages=[{"role": "user", "content": "Summarize this ticket"}],
)

Not ready to route production traffic?

Start with the usage you already have.

Import a usage export and see it in the real Cognocient dashboards, with waste flagged, before you change a line of code.

Import a usage CSV or OpenTelemetry export

Map your CSV columns in the importer, or upload OTel span JSON.

No SDK or code change No production access needed No credit card

Retry waste

A failed call retried shortly after on the same model with a near-identical prompt. The failed attempt's cost is flagged as waste.

Context bloat

Prompts that keep growing across a session, re-sending history you already paid for. Needs user and feature columns in your export.

Spend by feature and model

The same attribution dashboards the proxy fills, so you can see which features run on which models and what each costs.

Imported data is uploaded to and stored in your Cognocient account. It is historical only: nothing is monitored or enforced live until you route traffic through the proxy. Signing up here starts a 10-day trial with full history and waste detection; after it, the Free plan keeps the most recent 7 days.

Can't add a proxy hop? The Python wrapper gives live attribution with no added latency, but no pre-call enforcement.

Architecture

Where does Cognocient sit in your stack?

It replaces your gateway. Your tracing stays where it is.

Gateway & routing
A self-hosted gateway to run and patch, or direct SDK calls to each provider.
Replace it. One OpenAI-compatible endpoint for 10 provider types, with load balancing, the Auto Router, and budgets and loop breakers in the same hop. Nothing to host.
Tracing & prompt quality
Langfuse or another tracing tool records spans for review after the call.
Keep it. Cognocient decides before the call whether it is allowed, clamped, or moved to a cheaper model, and can export every call to Langfuse or any OpenTelemetry backend.
Financial accountability
Spreadsheets and a surprise invoice forwarded to Finance.
FOCUS-aligned chargeback exports and one-click board PDF reports by feature and team.

The waste taxonomy

Five places your AI budget disappears

None of these show up as a line item on a provider invoice. Four are flagged automatically, per call or in a daily scan; caching opportunities come from the Context Tax analysis.

01

Context Bloat

Multi-turn apps re-send the conversation on every call, so you pay again for history you already paid for. Flagged when a session's prompt grows more than 50% across its recent calls.

Per call · Base+
02

Retry Waste

A failed call gets retried within seconds with a near-identical prompt, often with no backoff. The failed attempt's cost is flagged as waste.

Per call · Base+
03

Model Mismatch

A premium model returning short answers without reasoning: the kind of task a smaller model often handles for a fraction of the price. The estimated saving is shown per call.

Per call · Base+
04

Cache Misses

The same large system prompt is sent fresh on every call, though providers discount cached input heavily. Context Tax finds features whose prompts barely change size between calls.

Context Tax · Growth+
05

Context Starvation

Consecutive calls from the same feature arrive seconds apart, each with more context than the last, because too little was sent up front and the same ground gets paid for twice.

Daily scan · Base+

Waste views are included on Base and above, and during the 10-day trial.

Platform capabilities

Everything your CFO and CTO have been asking for

Observe spend across every team and model. Enforce budgets in real time. Route to the right model automatically. Generate board-ready reports in one click.

10
provider types
2 min
setup time
1-click
recommendations

Illustrative previews of each view, with sample data, not customer results.

Every dollar, every team — in real time

Monthly spend by department — June 2026
Total Spend
$4,710
Waste Identified
$822
Budget Used
47%
Cost / Ticket
$0.42
Customer Supportgpt-4o-mini
$2,450
Engineering Toolsclaude-sonnet-4-6
$1,320
Sales Automationgpt-4o
$940
Updated every 5 min · 30-day windowExport as FOCUS CSV

See every platform feature in depth →

Ready for your security review

Route to any model. Control who can do what.

Bring AWS Bedrock, Vertex AI or your own servers into the same budgets and attribution, send each request to the right-sized model, balance load across models and keys, and give your team single sign-on, automatic provisioning, roles and an audit trail.

Auto Router

One model name; each request is classified by complexity and sent to the model you chose for that tier, with the reason and the saving on every response.

Growth and above

Bedrock, Vertex & self-hosted

AWS Bedrock, Google Vertex AI, and any OpenAI-compatible server such as Ollama or vLLM, all priced, budgeted and attributed like the rest.

Every plan

SAML single sign-on

Okta, Microsoft Entra ID, Google Workspace or any SAML IdP, with DNS-verified domains and an optional require-SSO policy.

Business

Roles & audit log

Owner, Admin, Developer, Finance and Viewer roles enforced on every call, and a searchable, exportable record of who changed what.

Roles on Base and up · audit log on Business

More endpoints, more routing

The Responses API, audio, rerank and the Batch API through one proxy, plus load balancing across models and providers and balancing across several API keys for one provider.

Endpoints on every plan · balancing on Growth and up

Agents, identity & secrets

An MCP gateway with tool policy for your agents, SCIM provisioning and JWT authentication from your identity provider, and provider keys kept in your own AWS, Vault or Azure vault.

Business

See every capability · Compare plans · Use cases

Every capability, included — hover a term to see what it does

Breaks down every AI API dollar by feature, team, or department — not just one lump monthly total.Enforce limits at run, feature, department, and org level at once — the tightest parent ceiling always wins.Groups related API calls into a session so you can see total cost per user story, ticket, or task.Automatic detection of cost spikes the moment they happen, with root-cause analysis attached.One-click PDF reports with an AI-written narrative, in Board/CFO, Finance, or Engineering tone.Measures what it actually costs to produce one resolved ticket, drafted contract, or business result — not just raw tokens.Links AI spend to the business outcomes it produced, so you can calculate real return on investment.30/60/90-day projections of where your AI bill is headed, built from your recent daily spend.Scores your AI FinOps practice from 0–100 using the Crawl/Walk/Run model, with next-action guidance.Flags features defaulting to an expensive frontier model for tasks a cheaper model handles equally well.Finds features paying a large, mostly-static system-prompt tax on every single call — prime caching candidates.Models how an agentic workflow's cost scales before you deploy it, so surprises show up in a simulation, not a bill.Exports spend data aligned with the FOCUS open standard for CloudZero, Apptio, AWS, and other FinOps tooling.Lets agents check a budget before spending, so a runaway loop hits a wall instead of your invoice.Sends each request to a model sized to its difficulty, using one model name and no code changes per call.AWS Bedrock and Google Vertex AI as first-class providers, priced and budgeted like the rest.Add any OpenAI-compatible server (Ollama, vLLM, LM Studio) as a custom endpoint, tracked alongside your cloud spend.Sign in with Okta, Microsoft Entra ID, Google Workspace or any SAML identity provider, with verified domains and optional enforcement.Owner, Admin, Developer, Finance and Viewer roles enforced on every API call, with least privilege by default.A searchable, exportable record of who changed what and when, including denied attempts. Request contents are never stored.Scrape month-to-date AI spend, tokens and budget utilisation into Prometheus or Grafana.OpenAI's Responses API on every provider, with streaming and function tools, under the same budgets and guardrails.Transcription, speech, rerank and the OpenAI Batch API through the proxy, costed and attributed like chat.One model name spread across models or providers by weight, rotation or lowest cost, skipping failing deployments.Several API keys per provider, balanced by weight and rate-limit aware, so one key's limit stops being your ceiling.Send every call, with cost, tokens, latency and your tags, to Langfuse or any OTLP backend.Govern MCP servers with tool allow and deny policy, argument scanning, per-call pricing and attribution.Provision users automatically from your identity provider and authenticate with its short-lived tokens instead of long-lived keys.Keep provider keys in AWS Secrets Manager, HashiCorp Vault or Azure Key Vault; fetched at request time, never stored.Fails a CI build when your test run's AI calls cost more than the baseline, so a cost regression is caught in the pull request. Base plan and above.Splits spend into production, staging, CI and evals with one header, and budgets each environment separately. Every plan.Compares prompt versions or models side by side on real cost per call, tokens, latency and errors. Growth plan and above.Ask Claude Code, Cursor or any MCP client about your spend, budgets and CI results. Read-only. Base plan and above.Custom charts from any metric, split and filter, shared with the team and pinned to the dashboard. Growth plan and above.Shadow Spend pulls the actual daily bill from OpenAI and Anthropic and flags spend the proxy never saw.

Works with every major AI provider

OpenAIGPT-4o · o3 · o4-mini
AnthropicClaude Sonnet · Opus · Haiku
Google Gemini2.5 · 3.x Pro and Flash
Azure OpenAIGPT-4o via Azure
AWS BedrockClaude · Nova · Llama
Google Vertex AIGemini on GCP
Mistral AILarge · Small
GroqLLaMA 3 · 70B
Together AIMeta LLaMA
Custom endpointOllama · vLLM · LM Studio

Building in the open

Cognocient is early.

We're working directly with our first customers to refine what CFOs and engineering teams actually need. Real case studies are coming as our first users get results.

Want to be one of the first?

Start your free trial

and tell us what you find.

Before you put us in your request path

Questions engineers ask first

How much latency does Cognocient add to my AI calls?

Our published benchmark (methodology and raw numbers at /docs/latency-benchmark) measured roughly 2-8ms of added latency at the median with 1-10 simultaneous requests to one instance: authentication, attribution and the atomic pre-call budget reservation. At high concurrency to a single instance (50-100 simultaneous requests) the overhead grows to hundreds of milliseconds, because the backend currently runs as a single worker process. That number is in the benchmark too, not hidden.

What happens if Cognocient goes down?

Two different cases. If Cognocient's budget store is unreachable, the proxy fails open by default: your calls pass through to the provider and you temporarily lose enforcement and attribution for that window. Individual budgets can be set to fail closed instead. If the Cognocient proxy itself is unreachable, calls sent to it fail like any other API outage, so keep a fallback to your provider's own URL in your client for critical paths. If you can't accept that dependency, the Python wrapper reports usage after each call and never sits in the request path, at the cost of pre-call enforcement.

Does Cognocient store my prompts or responses?

No. Cognocient stores call metadata (token counts, model, attribution tags, cost, latency), not prompt or response content. The one exception is response caching, which you opt into per request: a cached response is stored so it can be replayed. Provider API keys are encrypted at rest, and your application only ever holds a revocable Cognocient proxy key.

How does Cognocient work?

Point your OpenAI SDK (or any OpenAI-compatible client) at https://api.cognocient.com/v1 with a Cognocient proxy key. That base URL is the only code change; Claude, Gemini, Bedrock and Vertex models are reached through the same OpenAI format. Optionally add headers such as X-Cost-Feature and X-Cost-Session to attribute spend. Two lower-commitment paths also exist: a Python wrapper for live attribution without a proxy hop (no pre-call enforcement), and a CSV/OpenTelemetry importer that fills the dashboard from usage you already have (historical only).

How much does Cognocient cost?

There is a permanent Free plan (one provider, 7-day retention, no credit card) and three paid plans: Base at $99/mo, Growth at $499/mo and Business at $1,299/mo. Every new account can start a 10-day trial with every feature and no credit card; afterwards it stays on Free unless you pick a paid plan. Budget enforcement, Emergency Freeze and velocity limits are on every plan, including Free. See /pricing for the full comparison.

What is Cognocient?

An AI FinOps and agent-guardrail platform. It proxies the LLM API calls your application makes, to OpenAI, Anthropic, Google and other providers, attributes every dollar to a feature, team, user or session, and enforces budgets and loop limits before a call reaches the provider. Finance teams get chargeback exports and board-ready PDF reports from the same data.

How does Cognocient stop a runaway agent?

Several independent checks run before each call is forwarded. Per-run budgets shared across every call in a run (via X-Cost-Run-ID) cap what a whole agent run can spend, even when each call looks cheap. The Failure Loop Breaker (Growth and up) stops a session after 3 identical tool calls in a row or a run of consecutive errors. Dollar- and token-per-minute velocity limits catch fast floods. When a budget runs out, calls whose tools look like writes are hard-stopped while read-only calls can continue on a cheaper model. A handoff depth limit, configurable per budget, stops runaway agent-to-agent chains.

What types of AI waste does Cognocient detect?

Four are flagged automatically on Base and up: retry waste (a failed call retried within seconds with a near-identical prompt), model mismatch (a premium model returning short answers without reasoning), context bloat (a session's prompt growing more than 50% across its recent calls) and context starvation (a daily scan for consecutive calls that escalate in size). Caching opportunities come from the Context Tax analysis on Growth and up, which finds features whose prompts barely change size between calls.

How is Cognocient different from Langfuse, Helicone, or a self-hosted LLM gateway?

Tracing tools like Langfuse record what happened after the call, for engineers debugging prompts. A self-hosted gateway gives you routing and basic budget limits, and means running your own Redis, PostgreSQL and on-call rotation. Cognocient is a managed gateway that adds pre-call enforcement built for agents (loop breaking, per-run budgets, read/write-aware degradation) plus the finance layer: chargeback, board-ready PDF reports, cost per business outcome and FOCUS-aligned export. It can also export every call to Langfuse or any OpenTelemetry backend, so it sits alongside your tracing rather than replacing it.

Why does Cognocient use a proxy instead of read-only API access?

Tools that poll your provider's billing with a read-only key need no code change, but they can only report spend after the fact. Stopping a call before the provider bills you, whether by blocking it, degrading it to a cheaper model or breaking a loop, requires sitting in the request path. The trade-off is a one-line code change and a small measured latency cost (see the latency question above). If you only need after-the-fact reporting, a read-only tool may be the simpler fit.

Which AI providers does Cognocient support?

10 provider types: OpenAI, Anthropic, Google Gemini, Azure OpenAI, Mistral AI, Groq, Together AI, AWS Bedrock, Google Vertex AI, and custom OpenAI-compatible endpoints such as Ollama or vLLM (public HTTPS URLs only).

Can I put a hard cap on all my AI spend, not just a monthly budget?

Yes. Emergency Freeze is a one-click kill switch on every plan that blocks all proxied calls immediately, regardless of any budget. Freeze and unfreeze are self-service, and it can optionally trigger automatically on a spend-velocity spike or a high-severity compromise signal if you opt in; that is off by default.

Does Cognocient catch a compromised API key or account, not just cost spikes?

Two features target this. Compromise Risk (Base and up) flags a proxy key suddenly called from a new source IP or model family it has never used. Shadow Spend Reconciliation (Growth and up) compares what your provider actually billed against what the proxy observed, which catches spend that bypassed the proxy entirely, such as a raw API key called directly. For OpenAI and Anthropic it pulls the provider's daily cost live with an admin key; other providers use manual entry or a CSV upload.

What is in the board-ready reports?

On Base and up you can generate an AI ROI Summary or a Budget vs Actuals report on demand, as a PDF or an interactive view, with a 150-200 word narrative in Board/CFO, Technical Finance or Engineering tone. The narrative is written by the AI provider you have connected, using your own key. On Growth and up, the AI ROI Summary can be emailed on a schedule (monthly, weekly or on the last day of the month), optionally after an approver signs off.

Does Cognocient support single sign-on and role-based access control?

The Business plan includes SAML 2.0 single sign-on (Okta, Microsoft Entra ID, Google Workspace, or any SAML identity provider) with DNS-verified domains and an optional require-SSO policy, plus a searchable, exportable audit log. Owner, Admin, Developer, Finance and Viewer roles are enforced on every API call on every plan with team seats (Base and up).

Does Cognocient support image generation, like DALL-E or Imagen?

Yes. /v1/images/generations goes through the proxy like chat and embeddings, with the same budget, velocity and attribution coverage. OpenAI (DALL-E, gpt-image-1) and Google Imagen are supported and priced per image.

How do I know my own budget or velocity limit is not silently blocking legitimate calls?

Cognocient tracks your rejection rate as its own anomaly type, separate from spend and call-frequency spikes, with a 14-day baseline and a breakdown by reason. That catches both an attack that gets mostly blocked (so spend barely moves) and a limit set tight enough to reject real traffic.

Can Cognocient catch a cost regression before it ships?

Yes, on Base and above. Run your test suite through the Cognocient proxy with a run id, then one API call from CI compares that run with a baseline and fails the build if cost went up more than you allow (total, per call, per feature, or over an absolute cap). A GitHub Actions snippet is in the docs at /docs/ci-cost-gate.

Does Cognocient optimize GPU or cloud infrastructure costs?

No. Cognocient covers AI API spend billed by token or per request. It does not optimize GPU instances, cloud compute or Kubernetes; pair it with your cloud provider's tools for that.

Get started today

Protect your next agent deployment
in under two minutes.

Point your base URL at Cognocient, or import a usage export to see where your spend is going first. Every feature free for 10 days, no credit card.

No credit card No sales call Free forever on 1 provider after the trial Cancel any time in Settings → Billing Setup in 2 minutes