KONST
AI Training CalculatorContact us
EN
SERVICE | Compute data center operations

No need for your own operations team,
and compute runs 24/7 without a break

We take over day-to-day operations after a data center and GPU cluster go live: proactive alert monitoring, fault diagnosis, manufacturer RMA coordination, UFM cluster network upkeep, and emergency on-site response. Power and network outages get a one-hour response, equipment repair requests four hours, and a monthly operations report is produced. Sites we built and sites built by other vendors can both be commissioned.

Remote technical support24/7One-hour response to power and network outages
PROBLEM | Problems customers face

Delivering the equipment does not mean someone takes over daily operations

Transfer of ownership does not mean someone takes over the daily work. Alerts need dedicated on-duty staff, hardware faults need someone to open an RMA and chase it until the parts arrive, and the cluster network needs someone to isolate faulty nodes. All of this begins after go-live.

SOLUTION | KONST's approach

Daily operations cover six core tasks

On-duty staff watch alerts in real time and assess and classify them on the spot. Most incidents are resolved remotely, which gives the shortest response time. Once a hardware fault is confirmed, an RMA is opened and tracked until the parts arrive and the replacement is done. For incidents that need on-site work, staff are dispatched immediately and billed per visit, with rates tiered by weekday, holiday, and time of day.

  1. Problem

    No dedicated operations staff after go-live

    Monitoring, alerts, and RMA have no dedicated contact point, so existing IT staff handle them on top of their regular work.

    Solution

    Proactive monitoring and alert intake

    On-duty staff watch equipment status and alerts 24/7 in real time, and complete diagnosis, classification, and first-line resolution remotely.

    Solution

    AIOps intelligent operations platform

    All Day-2 remote operations run on this platform, which produces dashboards and reports, and security updates run on a schedule.

  2. Problem

    Skills gap in cluster networking

    GPU cluster networks differ greatly from ordinary data center networks, and without experience it is hard to locate faults promptly.

    Solution

    NVIDIA UFM cluster network operations

    Continuously monitors high-speed network status, detects link faults, and isolates faulty nodes.

    Solution

    Technical staff training

    The owner's technical staff can join training and take over first-line diagnosis and routine checks.

  3. Problem

    Gap between manufacturer warranty and on-site handling

    Warranty terms are written into the contract, but swapping parts, sending items for repair, and bringing machines back online still require someone on site.

    Problem

    No handling time limit agreed for fault resolution

    Without agreed response and on-site arrival times, you can only wait passively when a fault occurs.

    Solution

    Fault resolution and manufacturer RMA coordination

    After a hardware fault is confirmed, we open an RMA with the manufacturer and track it until the parts arrive and the replacement is done.

    Solution

    Emergency on-site support

    For incidents that need on-site work, an engineer comes on site. Billing is per visit, tiered by weekday or holiday and by whether the visit falls inside or outside service hours.

  1. 01Monitoring and alerts

    Continuously watch equipment and the cluster network.

  2. 02Assessment and classification

    Confirm the incident's impact and who is responsible for the response.

  3. 03Resolution and tracking

    Resolve the issue, coordinate the RMA, and keep records.

FEATURE | Service scope and specifications

Two service level modes,
chosen by scope of responsibility

When one party is responsible for both the equipment and the site, the availability guarantee mode applies: overall availability is the commitment, with an agreed service fee adjustment if the target is missed. When responsibility spans several parties, the response-time guarantee mode applies, committing to time limits for alert response and fault handling. Availability also depends on power, MEP, and equipment quality, so the mode and the responsibility boundary are confirmed together before signing.

Availability guarantee mode

Overall availability is the commitment. The service fee adjustment rate and cap that apply if the target is missed are stated in the contract.

Response-time guarantee mode

Commits to time limits for alert response and fault handling. Availability is assessed separately based on actual site conditions.

FEATURE | Service scope and specifications

The scope of the operations fee,
fully defined at signing

The boundary of the operations fee is written once at signing: which routine maintenance is included in the monthly fee, which parts and engineering work is priced separately, and compatibility issues after driver and framework upgrades.

Covered by the operations fee

The monthly fee covers the staffing and work of routine operations and is not priced separately by number of incidents.

Not included in the operations fee

Items that are parts costs, new engineering work, or the owner's own environment are quoted separately based on the actual items.

FEATURE | Service scope and specifications

Operations scope covers the whole compute path

Cluster faults can occur outside the GPU servers too: a switch link fault, a management node scheduling failure, or a storage mount interruption can each stop training. Operations must cover the whole path, and pricing is calculated from the unit count and type of each device.

  1. GPU servers

    Hardware alert monitoring, fault diagnosis, manufacturer RMA issuance and tracking

  2. Switches

    Link status monitoring, faulty port isolation, firmware version control

  3. Management and storage nodes

    Monitoring of scheduler service availability, mount status, and capacity levels

  4. Load balancers

    Topology health and cross-node bandwidth degradation detection

EVIDENCE | Track record and data

24/7 on-duty coverage

After go-live, the on-duty team takes over equipment monitoring, alert assessment, and hardware repair requests remotely. The status, handling, and timeline of every incident are fully recorded, so responsibility boundaries are clear and auditable.

24/7Remote technical support
MonthlyOperations report
FAQ | Common questions

FAQ

When a GPU cluster has a problem, how long until someone responds?
It depends on the incident category and the service level mode. Power and network outages affect the whole project, so their response time takes priority over ordinary equipment repair requests. The actual time limits are agreed in the contract based on site conditions.
What equipment does the operations contract cover?
Besides GPU servers, switches, management nodes, storage, load balancers, and firewalls must be included. A significant share of cluster outages come from the network layer rather than the compute nodes, so covering only the GPUs leaves a gap. The equipment list is confirmed before signing, and equipment added later is charged per unit.
How is the operations fee calculated?
It is calculated by the number and type of equipment. Before signing, an equipment list itemized unit by unit must be provided as the pricing basis, and equipment added later is charged per unit.
Can a short-term project sign for only a few months?
Contracts are annual, payment is a monthly fee, and invoices are issued month by month. For the transition period between build completion and the owner's team taking over, handover conditions can be agreed in the contract.
What reports are provided during operations?
A service record is produced every week. A monthly operations report includes alert statistics, fault handling records, RMA progress, and power and temperature trends.
We have our own IT team. Can we run operations together?
Yes. The owner's technical staff can join training and take over first-line diagnosis and routine checks. Academic and research institutions and companies with in-house IT often use this model.
What information security standards do KONST's data centers meet?
The data centers KONST operates itself or on behalf of others are certified to ISO 27001:2022. They have round-the-clock operations support, track incidents in a ticketing system, and commit to service levels under SLAs.

Is your operations team ready now that the data center is live?

Send us your equipment list, and we will reply with the operations scope and service level.