One project key, 70 models
ATP Token is an enterprise AI model management platform developed in-house by Horizon AI and offered as Token as a Service. With one project key, an enterprise connects to the 70 models and 5 modalities in the platform catalog. The interface is compatible with the SDKs of major providers, and in most cases you only change base_url and the key. Usage is metered per request and attributed to the organization, workspace, and project. Permissions are scoped to the key, and when a single upstream has a problem, the routing layer hands over to another source. You top up first, are charged as you use, and pay no monthly fee.
Four fragmentation risks of connecting to multiple AI models yourself
You maintain several sets of connection logic, AI cost cannot enter the allocation process, keys have no boundary, and you must add routing and failover yourself. None of this shows in the demo stage, but once you go live and external users appear, it becomes a hidden risk.
SDK, billing, keys, and failover converge on one governance plane
One project key connects to every model in the platform catalog. Usage is attributed per request, permissions are scoped to the key, and the routing layer keeps availability across multiple upstreams. An enterprise manages one platform, one bill, and one permission system.
- Problem
Scattered models: switching models means changing code
Each provider has its own SDK, keys, and model list, so switching models means changing code. Models iterate far faster than applications get revised.
Solution01 · ONE KEYUnified SDK
The interface is compatible with the SDKs of major providers, and requests and responses keep their original format. When migrating from OpenAI, Anthropic, or Google, in most cases you only change base_url and the key.
- Problem
Scattered billing: cost never reaches the allocation process
Each project's token usage and cost sit in separate provider consoles and cannot be combined into one report, so departmental allocation can only be estimated afterward.
Solution02 · ONE BILLUnified billing
Each project's token usage and cost are gathered into one report that feeds directly into the allocation process. Quota is allocated down the organization hierarchy, and an alert fires before a limit is exceeded.
- Problem
Scattered keys: no boundary on
permissions or usage Keys are spread across each project's environment variables, permission boundaries are unclear, and usage has no cap. Once external users appear, this becomes a risk.
Solution03 · ONE BOUNDARYUnified keys
Model authorization is set on the project, not on the key. Each key can call only the authorized models, usage caps and permission scope are set key by key, and logs are kept after revocation.
- Problem
Scattered availability: you add failover yourself
When a single provider has a fault, applies rate limits, or retires a model, you have to write the switching and retry logic yourself, and the risk of service interruption falls on the application.
Solution04 · ONE GATEWAYUnified failover
The routing layer picks a source among multiple upstreams based on availability. On an error or timeout, the next source takes over, with no manual switching in the application.
A request passes through four checkpoints inside the platform
Key verification, model authorization, upstream routing, and metering and charging run in that order. The first two checkpoints block requests inside the platform, and those requests are not sent to the upstream provider.
POST {ATP_GATEWAY}/v1/chat/completions
Authorization: Bearer <PROJECT_API_KEY>
{ "model": "<catalog-model-id>", "messages": [ ... ] }| Checkpoint | What the platform does | If it fails |
|---|---|---|
| 01 · AUTH Key verification | Verifies the project key. Upstream providers never see the ATP key. | Rejected if invalid (401) |
| 02 · SCOPE Model authorization | Checks whether the model is on the project's allowlist. | Rejected if unauthorized (403) |
| 03 · ROUTE Upstream routing | Selects a source from the supply pool and switches to the next source on an error or timeout. | The next source takes over |
| 04 · METER Metering and charging | Deducts from the project quota based on token usage. | Rejected if quota is insufficient (402) |
One plane, four roles, each solving one problem
With access, metering, and governance in one entry point, the benefits go beyond engineering. What finance and audit receive is what used to be pieced together after the fact.
Switching models is a setting change, not a new release
Switching models needs no new release: the interface is compatible with major SDKs, and in most cases migrating takes only two setting changes (service URL and key). New models are listed as soon as they launch, engineers do not need to apply for keys and integrate provider by provider, and the platform handles retries and failover.
Permissions converge, with no reliance on self-discipline
A four-level structure of organization → workspace → project → key, with model authorization set at the project level. Self-provisioned keys scattered across teams are retired and replaced by project keys under unified governance, which also narrows the Shadow AI gap.
Cost attributed to every request
Quota is deducted by input + output token counts multiplied by each model's rate, and every deduction maps to a project and feature. Quota is allocated down the hierarchy and usage is attributed request by request, so AI spending shifts from an end-of-month result to a live dashboard.
Compliance evidence is built into the platform
Each request records the model, key, token usage, and status. After a key is revoked, its logs remain in the audit record, so a security review does not need documents supplied afterward.
Same job: connecting yourself versus connecting through a governance plane
The table below compares approaches, not providers. Model capability is the same on both sides. The difference is who takes on access, billing, permissions, and failover.
| Aspect | Connect to each provider yourself | Govern with ATP Token |
|---|---|---|
| Switching models | Each provider has a different SDK and model list, so switching means changing code | Interface compatible with major SDKs; in most cases only base_url and the key change |
| Billing | Each provider sends its own bill, and you reconcile at month end | One contract, one bill, with usage attributed to projects request by request |
| Permissions | Keys scattered across environment variables, with scope left to developers' self-discipline | Model authorization is set at the project level, and a key can call only authorized models |
| Budget control | Overspend is found only when the bill arrives | Quota is allocated down the organization hierarchy, with an alert before the limit is reached |
| Availability | Switching and retry logic implemented and maintained yourself | The routing layer switches among multiple upstreams, and another source takes over on an error or timeout |
| Audit | Records are spread across consoles and have to be compiled afterward | Per-request logs keep the model, key, usage, and status, and remain after a key is revoked |
| Procurement | Negotiate and integrate provider by provider | Platform-wide contract pricing tailored to usage, so procurement and finance review happen once |
Five steps to bring the whole organization's model usage into one entry point
With your own account, you can complete the first three steps yourself. Enterprise-scale consolidation (inventory of existing keys, department quota design, contracts and billing) is supported by the adoption team.
- 01Create the organization and projectsConsole
Create workspaces and projects by department or application, invite members, and assign roles.
- 02Authorize models and allocate quotaConsole
Enable an allowed model list for each project, allocate budget quota down the organization hierarchy, and set caps.
- 03Change base_url and the keyApplication side
Existing applications do not need a new protocol. They keep their SDK and request format and connect to a single integration point.
- 04Call, route, and meterGo-live
Each request goes through key verification, model authorization, upstream routing, and quota deduction in order, and usage is recorded immediately.
- 05Audit and reconciliationOngoing
Use request logs to trace abnormal usage and the attribution report to complete departmental allocation. Once usage reaches scale, upgrade to an enterprise contract.
Model access, metering, and governance, brought into one entry point
The platform catalog and pricing are public information. Governance is first proven on Horizon AI's own projects, then offered to customers.
How we govern our own AI usage
All company model calls converge on ATP Token: one key, one quota, and one set of logs per customer project. The migration process is the standard adoption process we run for customers: inventory, consolidate, set quotas, and put dashboards in place, finished in two weeks.
- 70+models consolidated to a single integration point
- 100%of requests have auditable logs
- 0project quota overruns
Turning wasted token usage into a predictable cost
An internal enterprise AI agent running on multiple models uses project-level governance to bring scattered usage into a single billing point. Mismatched models and idle keys can be brought under control as soon as they are flagged.
- 100%of model calls consolidated into a single governance gateway
- 1consolidated bill replacing per-provider reconciliation
- 0token spend with no budget owner
Credit billing: top up first, pay as you use
There are no seat fees and no license seats. Pricing follows actual token usage, and you are charged for what you use. Once usage reaches scale, an enterprise contract can replace provider-by-provider negotiation.
| Item | Description |
|---|---|
| Billing method | Credits are deducted by input + output tokens multiplied by each model's rate, charged as each request completes. |
| Credits and top-up | 1 USD = 100 credits, with a minimum top-up of USD 5. Credits never expire, and there is no monthly fee or per-seat license fee. |
| Visibility | Usage and spend are visible in real time and attributed to organizations, workspaces, and projects. An alert fires before quota reaches its limit. |
| Enterprise contract | A platform-wide contract price tailored to actual usage: one contract, one bill, no provider-by-provider negotiation or integration, plus a dedicated solutions engineer to help with integration and tuning. |
| Term commitment | No term commitment by default. Enterprises can choose whether to sign a contract. |
| Adoption assessment | Qualifying enterprise applicants can receive proof-of-concept credits and a benchmark report, so they can estimate cost during the assessment phase. |
Common questions about access, billing, availability, and governance
What is Token as a Service, and how is it different from applying to each provider directly?
Token as a Service packages model access, metering, and governance as one service. An enterprise gets one project key from the platform and can call any model in the catalog. Usage is metered in tokens and attributed per request, and permissions are scoped to the key.
The difference from applying to providers one by one lies in the governance plane, not in the models. Model capability still comes from upstream providers. The platform brings access, spending, and auditing into one report and one permission system.
How much code has to change to move an existing application to the platform?
The interface is compatible with the SDKs of major providers, and requests and responses keep their original format. In most cases you only change base_url and the key. The platform sits on top of your existing call pattern and does not require the application to adopt a new protocol.
The list of available models can be queried with GET /v1/models. For actual endpoints and parameters, see the platform documentation.
If a single provider has a problem, are calls interrupted?
The routing layer switches among multiple upstream sources, and on an error or timeout another source takes over automatically, so the application does not need to switch manually.
Upstream providers never see the ATP key. Model authorization is checked at the project level, and a call to an unauthorized model is not sent upstream even if it is made.
How does the platform charge? Do credits expire?
It uses a credit system: input plus output tokens, multiplied by each model's rate, are deducted, and each request is charged as soon as it completes. 1 USD = 100 credits, with a minimum top-up of USD 5.
You top up first and are charged as you use, with no monthly fee and credits that never expire, and there is no per-seat license fee. Enterprises with larger usage can ask about enterprise contract pricing.
How is usage cost allocated to each department?
The platform uses a four-level structure: organization → workspace → project → key. Quota is allocated from the top down and usage is attributed from the bottom up, and each level shows available, allocated, and consumed quota.
The token count and charge of every call trace back to a project and feature, so each department's cost comes straight from the platform with no separate estimate.
How do you stop developers from running tasks on models that cost too much?
Model authorization is set on the project and does not depend on developers' self-discipline. Each project enables only the models its task tier needs, and that project's keys can call only those models, so lightweight tasks cannot reach flagship models.
With a project quota cap and request logs, mismatched models and idle keys surface and get fixed within the week, without waiting for the end-of-month bill.
After a key is revoked, can past usage records still be retrieved?
Yes. Each request records the model, key, token usage, and status. After a key is revoked, its logs stay in the audit record, so compliance reviews do not need evidence added afterward.
Can I use ATP Token as a platform only, without adoption services?
Yes. ATP Token is a platform service that can be adopted on its own. Create an organization and projects yourself and start using it.
If you need an inventory of existing keys, department quotas designed, ERP/CRM integration, or a custom application, Horizon AI's adoption team takes it on. Learn about enterprise AI adoption →
Consolidate scattered model access into one key
Start with an inventory of your current keys and usage, and we help plan the organization hierarchy, quotas, and permission boundaries.
