KONST
AI Training CalculatorContact us
EN
TOKEN AS A SERVICE

One project key, 70 models

ATP Token is an enterprise AI model management platform developed in-house by Horizon AI and offered as Token as a Service. With one project key, an enterprise connects to the 70 models and 5 modalities in the platform catalog. The interface is compatible with the SDKs of major providers, and in most cases you only change base_url and the key. Usage is metered per request and attributed to the organization, workspace, and project. Permissions are scoped to the key, and when a single upstream has a problem, the routing layer hands over to another source. You top up first, are charged as you use, and pay no monthly fee.

Models in the public catalog, with more being added70+models
PROBLEM | Problems customers face

Four fragmentation risks of connecting to multiple AI models yourself

You maintain several sets of connection logic, AI cost cannot enter the allocation process, keys have no boundary, and you must add routing and failover yourself. None of this shows in the demo stage, but once you go live and external users appear, it becomes a hidden risk.

SOLUTION | How ATP Token solves it

SDK, billing, keys, and failover converge on one governance plane

One project key connects to every model in the platform catalog. Usage is attributed per request, permissions are scoped to the key, and the routing layer keeps availability across multiple upstreams. An enterprise manages one platform, one bill, and one permission system.

  1. Problem

    Scattered models: switching models means changing code

    Each provider has its own SDK, keys, and model list, so switching models means changing code. Models iterate far faster than applications get revised.

    Solution01 · ONE KEY

    Unified SDK

    The interface is compatible with the SDKs of major providers, and requests and responses keep their original format. When migrating from OpenAI, Anthropic, or Google, in most cases you only change base_url and the key.

  2. Problem

    Scattered billing: cost never reaches the allocation process

    Each project's token usage and cost sit in separate provider consoles and cannot be combined into one report, so departmental allocation can only be estimated afterward.

    Solution02 · ONE BILL

    Unified billing

    Each project's token usage and cost are gathered into one report that feeds directly into the allocation process. Quota is allocated down the organization hierarchy, and an alert fires before a limit is exceeded.

  3. Problem

    Scattered keys: no boundary on permissions or usage

    Keys are spread across each project's environment variables, permission boundaries are unclear, and usage has no cap. Once external users appear, this becomes a risk.

    Solution03 · ONE BOUNDARY

    Unified keys

    Model authorization is set on the project, not on the key. Each key can call only the authorized models, usage caps and permission scope are set key by key, and logs are kept after revocation.

  4. Problem

    Scattered availability: you add failover yourself

    When a single provider has a fault, applies rate limits, or retires a model, you have to write the switching and retry logic yourself, and the risk of service interruption falls on the application.

    Solution04 · ONE GATEWAY

    Unified failover

    The routing layer picks a source among multiple upstreams based on availability. On an error or timeout, the next source takes over, with no manual switching in the application.

FEATURE | How the platform handles each request

A request passes through four checkpoints inside the platform

Key verification, model authorization, upstream routing, and metering and charging run in that order. The first two checkpoints block requests inside the platform, and those requests are not sent to the upstream provider.

REQUEST PATH
POST {ATP_GATEWAY}/v1/chat/completions
Authorization: Bearer <PROJECT_API_KEY>

{ "model": "<catalog-model-id>", "messages": [ ... ] }
CheckpointWhat the platform doesIf it fails
01 · AUTH Key verificationVerifies the project key. Upstream providers never see the ATP key.Rejected if invalid (401)
02 · SCOPE Model authorizationChecks whether the model is on the project's allowlist.Rejected if unauthorized (403)
03 · ROUTE Upstream routingSelects a source from the supply pool and switches to the next source on an error or timeout.The next source takes over
04 · METER Metering and chargingDeducts from the project quota based on token usage.Rejected if quota is insufficient (402)

Illustrative flow, taken from the platform's public documentation. For actual endpoints, parameters, and model IDs, refer to the atptoken.ai documentation.

BENEFIT | Benefits

One plane, four roles, each solving one problem

With access, metering, and governance in one entry point, the benefits go beyond engineering. What finance and audit receive is what used to be pieced together after the fact.

Development team

Switching models is a setting change, not a new release

Switching models needs no new release: the interface is compatible with major SDKs, and in most cases migrating takes only two setting changes (service URL and key). New models are listed as soon as they launch, engineers do not need to apply for keys and integrate provider by provider, and the platform handles retries and failover.

IT department

Permissions converge, with no reliance on self-discipline

A four-level structure of organization → workspace → project → key, with model authorization set at the project level. Self-provisioned keys scattered across teams are retired and replaced by project keys under unified governance, which also narrows the Shadow AI gap.

Finance

Cost attributed to every request

Quota is deducted by input + output token counts multiplied by each model's rate, and every deduction maps to a project and feature. Quota is allocated down the hierarchy and usage is attributed request by request, so AI spending shifts from an end-of-month result to a live dashboard.

Audit and security

Compliance evidence is built into the platform

Each request records the model, key, token usage, and status. After a key is revoked, its logs remain in the audit record, so a security review does not need documents supplied afterward.

COMPARISON | Self-built access vs. platform governance

Same job: connecting yourself versus connecting through a governance plane

The table below compares approaches, not providers. Model capability is the same on both sides. The difference is who takes on access, billing, permissions, and failover.

AspectConnect to each provider yourselfGovern with ATP Token
Switching modelsEach provider has a different SDK and model list, so switching means changing codeInterface compatible with major SDKs; in most cases only base_url and the key change
BillingEach provider sends its own bill, and you reconcile at month endOne contract, one bill, with usage attributed to projects request by request
PermissionsKeys scattered across environment variables, with scope left to developers' self-disciplineModel authorization is set at the project level, and a key can call only authorized models
Budget controlOverspend is found only when the bill arrivesQuota is allocated down the organization hierarchy, with an alert before the limit is reached
AvailabilitySwitching and retry logic implemented and maintained yourselfThe routing layer switches among multiple upstreams, and another source takes over on an error or timeout
AuditRecords are spread across consoles and have to be compiled afterwardPer-request logs keep the model, key, usage, and status, and remain after a key is revoked
ProcurementNegotiate and integrate provider by providerPlatform-wide contract pricing tailored to usage, so procurement and finance review happen once
PROCESS | Workflow

Five steps to bring the whole organization's model usage into one entry point

With your own account, you can complete the first three steps yourself. Enterprise-scale consolidation (inventory of existing keys, department quota design, contracts and billing) is supported by the adoption team.

  1. 01Create the organization and projectsConsole

    Create workspaces and projects by department or application, invite members, and assign roles.

  2. 02Authorize models and allocate quotaConsole

    Enable an allowed model list for each project, allocate budget quota down the organization hierarchy, and set caps.

  3. 03Change base_url and the keyApplication side

    Existing applications do not need a new protocol. They keep their SDK and request format and connect to a single integration point.

  4. 04Call, route, and meterGo-live

    Each request goes through key verification, model authorization, upstream routing, and quota deduction in order, and usage is recorded immediately.

  5. 05Audit and reconciliationOngoing

    Use request logs to trace abnormal usage and the attribution report to complete departmental allocation. Once usage reaches scale, upgrade to an enterprise contract.

EVIDENCE | Track record and data

Model access, metering, and governance, brought into one entry point

The platform catalog and pricing are public information. Governance is first proven on Horizon AI's own projects, then offered to customers.

70+open-source modelsModels in the public catalog, with more being added
5modalitiesText, image, video, speech, and embeddings
11providersBrought into one contract and bill
Zeromonthly feeTop up first, pay as you use; credits do not expire
CASE · Horizon AI (in-house validation)

How we govern our own AI usage

All company model calls converge on ATP Token: one key, one quota, and one set of logs per customer project. The migration process is the standard adoption process we run for customers: inventory, consolidate, set quotas, and put dashboards in place, finished in two weeks.

  • 70+models consolidated to a single integration point
  • 100%of requests have auditable logs
  • 0project quota overruns
Read the full case study →
CASE · Taiwanese AI startup expanding internationally (anonymized)

Turning wasted token usage into a predictable cost

An internal enterprise AI agent running on multiple models uses project-level governance to bring scattered usage into a single billing point. Mismatched models and idle keys can be brought under control as soon as they are flagged.

  • 100%of model calls consolidated into a single governance gateway
  • 1consolidated bill replacing per-provider reconciliation
  • 0token spend with no budget owner
Read the full case study →

The number of models, modalities, and providers comes from the public atptoken.ai catalog and changes as the platform updates. Case figures are actual results from individual deployments and vary by project conditions. Model capability still comes from upstream providers, and ATP Token brings access, spending, and auditing onto one plane.

PRICE | Pricing and fees

Credit billing: top up first, pay as you use

There are no seat fees and no license seats. Pricing follows actual token usage, and you are charged for what you use. Once usage reaches scale, an enterprise contract can replace provider-by-provider negotiation.

ItemDescription
Billing methodCredits are deducted by input + output tokens multiplied by each model's rate, charged as each request completes.
Credits and top-up1 USD = 100 credits, with a minimum top-up of USD 5. Credits never expire, and there is no monthly fee or per-seat license fee.
VisibilityUsage and spend are visible in real time and attributed to organizations, workspaces, and projects. An alert fires before quota reaches its limit.
Enterprise contractA platform-wide contract price tailored to actual usage: one contract, one bill, no provider-by-provider negotiation or integration, plus a dedicated solutions engineer to help with integration and tuning.
Term commitmentNo term commitment by default. Enterprises can choose whether to sign a contract.
Adoption assessmentQualifying enterprise applicants can receive proof-of-concept credits and a benchmark report, so they can estimate cost during the assessment phase.

For each model's actual rates, see the atptoken.ai pricing page, which is updated as upstream providers adjust. The sales team confirms eligibility for enterprise contracts and proof-of-concept credits case by case.

FAQ | Common questions

Common questions about access, billing, availability, and governance

What is Token as a Service, and how is it different from applying to each provider directly?

Token as a Service packages model access, metering, and governance as one service. An enterprise gets one project key from the platform and can call any model in the catalog. Usage is metered in tokens and attributed per request, and permissions are scoped to the key.

The difference from applying to providers one by one lies in the governance plane, not in the models. Model capability still comes from upstream providers. The platform brings access, spending, and auditing into one report and one permission system.

How much code has to change to move an existing application to the platform?

The interface is compatible with the SDKs of major providers, and requests and responses keep their original format. In most cases you only change base_url and the key. The platform sits on top of your existing call pattern and does not require the application to adopt a new protocol.

The list of available models can be queried with GET /v1/models. For actual endpoints and parameters, see the platform documentation.

If a single provider has a problem, are calls interrupted?

The routing layer switches among multiple upstream sources, and on an error or timeout another source takes over automatically, so the application does not need to switch manually.

Upstream providers never see the ATP key. Model authorization is checked at the project level, and a call to an unauthorized model is not sent upstream even if it is made.

How does the platform charge? Do credits expire?

It uses a credit system: input plus output tokens, multiplied by each model's rate, are deducted, and each request is charged as soon as it completes. 1 USD = 100 credits, with a minimum top-up of USD 5.

You top up first and are charged as you use, with no monthly fee and credits that never expire, and there is no per-seat license fee. Enterprises with larger usage can ask about enterprise contract pricing.

How is usage cost allocated to each department?

The platform uses a four-level structure: organization → workspace → project → key. Quota is allocated from the top down and usage is attributed from the bottom up, and each level shows available, allocated, and consumed quota.

The token count and charge of every call trace back to a project and feature, so each department's cost comes straight from the platform with no separate estimate.

How do you stop developers from running tasks on models that cost too much?

Model authorization is set on the project and does not depend on developers' self-discipline. Each project enables only the models its task tier needs, and that project's keys can call only those models, so lightweight tasks cannot reach flagship models.

With a project quota cap and request logs, mismatched models and idle keys surface and get fixed within the week, without waiting for the end-of-month bill.

After a key is revoked, can past usage records still be retrieved?

Yes. Each request records the model, key, token usage, and status. After a key is revoked, its logs stay in the audit record, so compliance reviews do not need evidence added afterward.

Can I use ATP Token as a platform only, without adoption services?

Yes. ATP Token is a platform service that can be adopted on its own. Create an organization and projects yourself and start using it.

If you need an inventory of existing keys, department quotas designed, ERP/CRM integration, or a custom application, Horizon AI's adoption team takes it on. Learn about enterprise AI adoption →

Consolidate scattered model access into one key

Start with an inventory of your current keys and usage, and we help plan the organization hierarchy, quotas, and permission boundaries.