CCS AI Inference Gateway
An API for running AI models from your own code, operated by the Center for Computational Sciences. It is not a chat website. It speaks the OpenAI API and forwards each request straight to the model gateway, so any tool that already talks to OpenAI talks to this: models hosted on CCS hardware on every key, and commercial cloud models on a project key.
Every figure above is inference on CCS hardware, updated every half-minute. Commercial cloud models, in limited testing and not available for general use, have generated a further 100,328,017 tokens.
- ccs/Qwen/Qwen3.6-35B-A3B-FP8 230,752 req
- ccs/gpt-oss:120b 152,072 req
- ccs/qwen3:32b 141,318 req
- ccs/deepseek-r1:32b 48,590 req
- ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink 20,251 req
Available models
8 live from the upstream catalog · detail and healthBefore you use this service
Model output is not authoritative. These are statistical language models. They produce fluent text that can be wrong, incomplete, or fabricated, including citations, code and numbers. Verify anything you intend to rely on.
Do not submit restricted data. This service is approved for public and internal-use data only. Do not send HIPAA, FERPA, ITAR/EAR, CUI, personally identifiable information, human-subjects data, or anything else covered by a data use agreement or IRB protocol.
Your prompts and the models' responses are not stored. Neither this host nor the CCS model gateway keeps any copy of what you send to a model or what it answers. What is logged is request metadata only -- timestamp, model name, response status, token counts, latency and client address -- for operations and capacity planning. Requests served by commercial cloud models are additionally subject to the cloud provider's data terms.
No availability guarantee. This is a research service on shared University hardware. Models may be added, changed or withdrawn without notice, and the service may be interrupted for maintenance.
Use is subject to University of Kentucky acceptable use policy and the terms of the underlying model licences.
Three ways to run open-source AI models at CCS
This site is the managed LLM API (the medium option). Two others exist for different scales of work.
| Scale of work | Option | Runs on | What to do |
|---|---|---|---|
| Large | Run your own models on cluster GPUs | HPC cluster GPUs (LCC / MCC / ECC) | Self-Hosted LLMs on GPUs (batch / CLI) |
| Medium (here) | Managed LLM inference API | Shared NVIDIA Grace Hopper cluster (GH200) | Start here. Get a key and use it, see below |
| Small | DGX Spark (low-cost, slower) | NVIDIA DGX Spark units, in pairs | Open a support ticket |
DGX Spark has slower GPUs than the managed API and the HPC clusters, and runs a smaller set of models; CCS can also host DGX Sparks you buy (condo). Ask in a ticket.
Research service. Do not enter sensitive or regulated data. Acceptable use and data policy.
Getting started
-
Get a key
LLM access is granted per person and does not require a compute allocation. The steps are the same for everyone:
- Sign in to Locksmith with your university login. This creates your account and your Locksmith user id — do this first, because access is granted to an account that already exists.
- Ask your PI to request LLM access for your Locksmith user id at the CCS service desk. No compute allocation is required — the PI picks the form that fits the group:
- Group is new to CCS (no allocation yet): this request form.
- Group already has a CCS allocation: this request form.
- CCS staff review the request and email you once it is approved.
- Back in Locksmith, generate your personal API key (shown once) and use it with the base URL below.
-
Point your client here
Set the base URL and use your key as the API key. No other change: no special library, no adapter.
Base URLhttps://llm.ccs.uky.edu/v1Nothing installed yet? Try it in your browser. Paste the key, pick a model, see the answer, and copy the same call as curl or Python.
-
Pick a model
Model ids look like
ccs/llama3.1:8b. The status page lists every one that is live, what it is good at, and what it is not.Wiring up a tool, like a chat UI, an editor, a notebook, or a SLURM job? The documentation has the exact settings for each.