Everything between your apps and the model
17 capabilities across three jobs: secure the traffic, optimise it, and stay in control.
One API for every model
Point your existing SDK at the gateway and keep your code. The same endpoint reaches OpenAI, Anthropic, Azure AI Foundry, AWS Bedrock, Google Gemini, Mistral and self-hosted models, so you can change model or provider without a rewrite.
- /v1/chat/completions, /v1/responses, /v1/embeddings and /v1/messages, plus images, speech and transcription
- Works with the official OpenAI and Anthropic SDKs, plus any HTTP client
- Switch provider by changing a model name, not your application
- Self-hosted and private models sit behind the same controls
Virtual keys, teams and budgets
Issue a gateway key to each application, team or person. Each key carries its own allowed models, spend budget and rate limit, and can be revoked in a click without touching a provider account.
- Gateway virtual keys with expiry and revocation
- Teams and users, with per-key model permissions
- Budgets with alerts, and requests-per-minute rate limits
- Provider keys stay inside the gateway. Applications never hold them
PII protection, tuned for Australia
Sensitive data is detected before a prompt leaves your control. Choose what happens for each kind of data: watch it, mask it, swap it for a reversible token, or stop the request. Detection is probabilistic, so test it against your own data.
- Australian identifiers: Medicare number, TFN, ABN, driver licence, passport
- New Zealand identifiers: IRD number, NHI number and bank account numbers, with a dedicated NZ policy preset
- Also date of birth and card numbers (including test card numbers), plus names, emails, phone numbers and addresses
- Four modes: monitor, redact, tokenise (restored on response) or block
- Per-policy control over which entity types are acted on
Same prompt, four modes (illustrative)
Secrets and credential detection
Developers paste logs and config into chat. The gateway recognises common key formats, tokens and private key blocks and can block or mask them before they reach a third party.
- Cloud, source-control and payment API key patterns
- Private keys, bearer tokens and connection strings
- Block or redact, with an audit record of what was caught
- Applies to prompts and to model responses
Prompt-injection detection
Requests and retrieved content are checked for attempts to override your instructions or extract hidden prompts. Run in monitor mode first to see what would be caught, then enforce.
- Instruction-override and system-prompt extraction patterns
- Monitor first, enforce when you are confident
- Events logged with the key, team and model involved
- Layered defence: a control, not a guarantee against every attack
Response DLP scanning
Model output can contain sensitive data too, from your own context or from a connected tool. Responses are scanned with the same detectors and the same four modes.
- Same detectors and modes as requests
- Tokenised values restored on the way back to your application
- Streaming-aware handling
- Separate policy for requests and responses
Safe prompt optimisation
The gateway normalises whitespace, compacts JSON and removes duplicated context. You see a before and after token count for each change. It never strips punctuation blindly, because that changes what a prompt means.
- Whitespace, JSON and duplicate-context normalisation
- Before and after token savings shown per request
- Code, quoted text and structured data left intact
- On or off per policy
Optimisation report (illustrative)
Response caching
Identical requests can be answered from cache at no model cost and with lower latency. Semantic caching, which matches similar rather than identical prompts, is off until you switch it on, because it suits some workloads and not others.
- Exact-match response caching
- Opt-in semantic caching with a similarity threshold
- Cache keyed per policy so tenants never share answers
- Hits and estimated savings reported in the console
Smart model routing
Route simple work to faster, lower-cost models and keep demanding work on premium ones. Every routing decision includes a plain-language "why this model" explanation so you can trust it and tune it.
- Rules by task type, size, team or key
- "Why this model" explanation on every decision
- Never routes outside the models a key is allowed to use
- Easy to turn off for sensitive workloads
Why this model (illustrative)
Retries and failover
Transient errors are retried and, when a provider or model is unavailable, traffic fails over to the alternatives you have approved.
- Automatic retries with back-off
- Provider and model failover chains
- Failover only uses models the key is already permitted to use
- Failover events visible in logs
Spend tracking and alerts
Every request is attributed so finance and engineering can see where money is going. Set budgets, get alerts before limits are hit, and see estimated savings from caching, optimisation and routing.
- Breakdowns by app, team, user, model and provider
- Budgets, thresholds and alerts
- Estimated savings clearly labelled as estimates
- Export for chargeback and reporting
Audit logs and OpenTelemetry
Who called which model, under which policy, and what the gateway did about it. Administrative changes to organisations, keys, teams, members and invitations are written to an append-only, tamper-evident audit log that Tokard staff can review. Send traces and metrics to the observability stack you already run using OpenTelemetry.
- Append-only, tamper-evident log of administrative changes, visible to staff
- Content logging is configurable, so you can keep prompts out of logs
- OpenTelemetry export for traces and metrics
- Request and policy-action records attributed to key, team and model
Policy presets, every feature a switch
Eight presets give you a sensible starting point, and every individual feature is a simple on/off switch you can override per team or key.
- Default, Healthcare, Finance, Government, Developer, High Security, AI Agent and New Zealand presets
- Every feature is an on/off switch
- Assign a preset to a team, key or environment
- Preview in monitor mode before enforcing
Sign-in: Entra for staff, invitations for customers
Microsoft Entra single sign-on is used for the platform's own staff and admin console. Customers sign in at the customer login with the account their organisation administrator invites, using email and password.
- Staff and administrators sign in to the admin console with Microsoft Entra
- Customers sign in at /customerlogin with an emailed invitation, then email and password
- Role-based access: customers see only their own organisation
- Applications authenticate with gateway keys, not logins
More than chat
One OpenAI-compatible gateway handles chat, responses, embeddings, Anthropic-style messages, image generation, text-to-speech and speech-to-text transcription, all under the same keys, budgets and policies.
- PII guardrails apply to text-to-speech, embedding and image prompts
- Transcription: the returned transcript is scanned. Healthcare-style policies mask identifiers one way, finance and government-style policies refuse the request
- The audio itself reaches your chosen provider unredacted. Only the transcript text is scanned
- Same keys, budgets and model permissions for every endpoint
Organisations and customer login
Set up customer organisations with their own administrators, teams, keys and budgets. Customers sign in at the customer login and see only their own organisation's keys, usage and cost.
- Organisations with their own admins, teams, keys and budgets
- Customer login shows only that organisation's keys, usage and cost
- Administrators invite members by email
- Provider keys stay inside the gateway
USD and AUD cost display
Costs can be shown in US dollars or Australian dollars at the touch of a switch, using the live ECB exchange rate. API responses carry cost headers in both currencies, so your own systems can record either one.
- Switch every cost view between USD and AUD
- Live ECB exchange rate
- Cost headers in both currencies on API responses
- Exchange rates move, so converted amounts are indicative
Keep every AI interaction within bounds
Request access and we will help you set up your first policy, key and budget.