CCS AI Inference Gateway

An API for running AI models from your own code, operated by the Center for Computational Sciences. It is not a chat website. It speaks the OpenAI API and forwards each request straight to the model gateway, so any tool that already talks to OpenAI talks to this: models hosted on CCS hardware on every key, and commercial cloud models on a project key.

1,451,798
Tokens generated today
910 requests
123,678,136
Tokens generated, last 7 days
858,164,484
Tokens generated, last 30 days
3,274,490,225
Tokens generated, since April 2026

Every figure above is inference on CCS hardware, updated every half-minute. Commercial cloud models, in limited testing and not available for general use, have generated a further 100,328,017 tokens.

Busiest models
Operational
Upstream gateway
8
Models available
3d 3h
Broker uptime
Status
Per-model health and detail

Available models

8 live from the upstream catalog · detail and health

Before you use this service

Model output is not authoritative. These are statistical language models. They produce fluent text that can be wrong, incomplete, or fabricated, including citations, code and numbers. Verify anything you intend to rely on.

Do not submit restricted data. This service is approved for public and internal-use data only. Do not send HIPAA, FERPA, ITAR/EAR, CUI, personally identifiable information, human-subjects data, or anything else covered by a data use agreement or IRB protocol.

Your prompts and the models' responses are not stored. Neither this host nor the CCS model gateway keeps any copy of what you send to a model or what it answers. What is logged is request metadata only -- timestamp, model name, response status, token counts, latency and client address -- for operations and capacity planning. Requests served by commercial cloud models are additionally subject to the cloud provider's data terms.

No availability guarantee. This is a research service on shared University hardware. Models may be added, changed or withdrawn without notice, and the service may be interrupted for maintenance.

Use is subject to University of Kentucky acceptable use policy and the terms of the underlying model licences.

Three ways to run open-source AI models at CCS

This site is the managed LLM API (the medium option). Two others exist for different scales of work.

Scale of workOptionRuns onWhat to do
LargeRun your own models on cluster GPUsHPC cluster GPUs (LCC / MCC / ECC)Self-Hosted LLMs on GPUs (batch / CLI)
Medium (here)Managed LLM inference APIShared NVIDIA Grace Hopper cluster (GH200)Start here. Get a key and use it, see below
SmallDGX Spark (low-cost, slower)NVIDIA DGX Spark units, in pairsOpen a support ticket

DGX Spark has slower GPUs than the managed API and the HPC clusters, and runs a smaller set of models; CCS can also host DGX Sparks you buy (condo). Ask in a ticket.

Research service. Do not enter sensitive or regulated data. Acceptable use and data policy.

Getting started

  1. Get a key

    LLM access is granted per person and does not require a compute allocation. The steps are the same for everyone:

    1. Sign in to Locksmith with your university login. This creates your account and your Locksmith user id — do this first, because access is granted to an account that already exists.
    2. Ask your PI to request LLM access for your Locksmith user id at the CCS service desk. No compute allocation is required — the PI picks the form that fits the group:
    3. CCS staff review the request and email you once it is approved.
    4. Back in Locksmith, generate your personal API key (shown once) and use it with the base URL below.
  2. Point your client here

    Set the base URL and use your key as the API key. No other change: no special library, no adapter.

    Base URL https://llm.ccs.uky.edu/v1

    Nothing installed yet? Try it in your browser. Paste the key, pick a model, see the answer, and copy the same call as curl or Python.

  3. Pick a model

    Model ids look like ccs/llama3.1:8b. The status page lists every one that is live, what it is good at, and what it is not.

    Wiring up a tool, like a chat UI, an editor, a notebook, or a SLURM job? The documentation has the exact settings for each.