Budgets & Control

What happens if Cognocient's proxy goes down?

Two different failures: if Cognocient's budget-check layer is unreachable, calls fail open and pass through to your provider. If the proxy itself is unreachable, calls sent to it fail, so keep a fallback for critical paths.

"Cognocient going down" can mean two different things, and they behave differently:

  1. The budget-check layer is unreachable (the proxy is up, but the Redis store it checks budgets against is not). The proxy fails open: calls pass through to your AI provider within a fraction of a second. You temporarily lose cost visibility and enforcement, but your application keeps working. Most of this page is about this case.
  2. The proxy itself is unreachable (a Cognocient deploy, crash or network problem). Your application is calling Cognocient's URL, so those requests fail the same way they would if any API you depend on were down. Nothing on Cognocient's side can fail open here. See If the proxy itself is unreachable for how to protect critical paths.

This is a real, verified behavior — not a promise

This isn't aspirational copy. It's the actual behavior of the budget-check code path, confirmed by reading it directly: on a Redis error the check logs a warning and returns "allowed", rather than raising and blocking the request.

If the proxy itself is unreachable

When your application can't reach api.cognocient.com at all, there is no Cognocient code running to make a decision, so the call fails with a connection error or timeout. If that's unacceptable for a feature, you have two options:

  • Fall back to the provider in your client. Catch connection errors and timeouts from the Cognocient client and retry the same request against the provider directly. Calls made this way skip Cognocient entirely: no budget check, no attribution, and they will show up as unobserved spend in Shadow Spend.

    from openai import OpenAI, APIConnectionError, APITimeoutError
     
    cog = OpenAI(api_key=COGNOCIENT_KEY, base_url="https://api.cognocient.com/v1", timeout=30)
    direct = OpenAI(api_key=OPENAI_KEY)  # fallback only
     
    def chat(**kwargs):
        try:
            return cog.chat.completions.create(**kwargs)
        except (APIConnectionError, APITimeoutError):
            return direct.chat.completions.create(**kwargs)

    Only fall back on errors that mean Cognocient couldn't be reached. Don't fall back on a 429 or 403 from Cognocient: those are your budgets, velocity limits or Emergency Freeze doing their job.

  • Use the Python wrapper instead of the proxy. It calls your provider directly and reports usage to Cognocient afterwards, so a Cognocient outage never touches the request itself. The trade-off is that it can't enforce anything before a call is made.

Why fail-open is the right default

Blocking every production LLM call because a cost-tracking layer is temporarily down would trade a minor, recoverable problem — a gap in attribution data — for a major, unrecoverable one: your actual product failing for your users. For the overwhelming majority of teams, a few minutes of unattributed spend is a far smaller cost than a production outage in a customer-facing AI feature. Cognocient is built to protect your uptime first.

The one exception: if a block-mode budget's last known spend was already over its limit before the outage started, Cognocient keeps blocking that budget's calls using that last-known value rather than assuming it's now safe to let everything through blind. Fail-open means "don't newly block on missing data" — not "ignore data you already have."

What you lose during a fail-open window

  • Attribution data for calls made during the outage — cost, tokens, latency, and feature/team tags for those specific calls will be incomplete or delayed until the connection recovers.
  • Budget enforcement during the outage — a block-mode budget will not newly block calls it hasn't already confirmed are over limit. alert and degrade-mode budgets already never block calls, so nothing changes for those regardless of Redis health.

Nothing about this affects calls that complete successfully — you keep making API calls and getting responses. The gap is purely in cost visibility and enforcement for the duration of the outage.

Choosing fail closed for a specific budget

Fail open is the right default for the overwhelming majority of traffic, but not every workload has the same risk tolerance. Every budget — not the account as a whole — has its own If Cognocient's budget check is unreachable setting, defaulting to fail open. Switch it to fail closed and calls matching that budget are blocked (HTTP 503, type: "budget_backend_unavailable") rather than let through unmetered during an outage.

This is deliberately per-budget rather than a single account-wide toggle: a marketing chatbot budget and a budget scoped to a workload that executes financial transactions have different answers to "would I rather this call fail, or succeed without being checked?" Set it from the budget's create/edit form on the Budgets page.

Fail closed has a real cost

A fail-closed budget turns a Cognocient outage into a production outage for whatever traffic matches it. Only choose this for the specific budgets where an unmetered call is genuinely worse than a blocked one — not as a default.

Monitoring and alerting

Honest gap, not a feature we're hiding

There is currently no dedicated customer-facing alert that fires the moment a fail-open window starts. The event is logged server-side, but isn't yet surfaced as a notification. If you have block-mode budgets and want certainty about enforcement continuity, the most reliable signal today is a visible gap or delay in attribution data on your dashboard for a period you know had real traffic — check Live Calls for the affected window. We're tracking a dedicated fail-open alert as a roadmap item.

If you have a critical-path AI feature where a temporary enforcement gap is unacceptable regardless of likelihood, treat that as a reason to keep your own downstream rate limits or spend caps as a backstop — not a reason to avoid the proxy, since the alternative (fail-closed) would make a Cognocient outage a worse problem for you than the one it solves.

Frequently asked questions

On this page