Service status

Jump to models

Operational The upstream inference gateway is reachable.

8 responding Every model is probed with one minimal request every 15 minutes; the last sweep ran 12 min ago. “Responding” means the model answered that probe (it checks that the model replies, not that the answer is correct) and it is a recorded result, not a call made when you loaded this page.

Four NVIDIA GH200 Grace Hopper servers serve these models. Each pairs an H100-class GPU (96 GB HBM3) with a 72-core Grace CPU and ~480 GB of LPDDR5X - about 576 GB of unified memory over NVLink-C2C.

Models

8 live · notes updated 2026-08-17 · service status

Reasoning models spend tokens thinking before they answer

A reasoning model uses part of your max_tokens budget on internal reasoning. If the budget is small it can spend ALL of it thinking and return an empty string -- which looks like a broken service but is not. Measured here: ccs/qwen3:32b used 289 completion tokens to answer "What is 2+2". At max_tokens=24 the same question returned "". Ask reasoning models for at least 2,000 tokens, or use a non-thinking model such as ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink or ccs/llama3.1:8b.

chat
General instruction-following. Ask it something, get an answer.
reasoning
Thinks step by step before answering. Better at hard problems, slower, and it spends tokens on the thinking.
vision
Accepts images alongside text.
embeddings
Turns text into vectors for search and RAG. Does NOT answer questions.

Qwen3.6 35B-A3B (no-think)

Responding
chat

Same model, thinking switched off. The fastest thing here.

ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink
Size35B / 3B per token active ArchMoE Contextabout 200 pages (approximately 131,000 tokens)
no thinkingtools
Good for
  • Interactive chat and autocomplete where latency is what you feel
  • Claude Code and similar tools that dislike a thinking preamble
  • Short, well-specified tasks that do not need deliberation
Detail, limits and gotchas
Notes
  • Measured fastest of all eight: about 0.10s to a short non-streaming reply.
  • Not a different model. If you want deliberation, use the thinking alias instead.
Deployment
  • Alias of ccs/Qwen/Qwen3.6-35B-A3B-FP8, same weights, same GPU, reasoning disabled.
  • Limit is about 200 pages (approximately 131,000 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.

Upstream model card

Last check 12 min agoLatency 37msUptime 24h 100.0% (96 checks)

Llama 3.1 8B

Responding
chat

Small, quick, dependable. Good default for simple work.

ccs/llama3.1:8b
Size8B ArchDense Contextabout 25 pages (approximately 15,500 tokens)
tools
Good for
  • Classification, extraction, summarising, rewriting
  • High-volume jobs where per-request latency matters
  • A cheap first pass before escalating to a bigger model
Detail, limits and gotchas
Watch out
  • This is NOT a search / embedding model, but it will still accept a 'turn this text into search numbers' request and hand back 4,096 numbers that look real and are not. A search tool built on them quietly returns poor results, and nothing in the reply reveals the problem. Send embedding and semantic-search requests only to ccs/Qwen/Qwen3-Embedding-8B.
Notes
  • No thinking phase, so what you set as max_tokens is what you get as answer.
Deployment
  • Limit is about 25 pages (approximately 15,500 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.

Upstream model card

Last check 12 min agoLatency 107msUptime 24h 100.0% (96 checks)

Qwen3.6 35B-A3B

Responding
reasoning

The default pick. Big-model quality at small-model speed.

ccs/Qwen/Qwen3.6-35B-A3B-FP8
Size35B / 3B per token active ArchMoE Contextabout 200 pages (approximately 131,000 tokens)
thinks firsttools
Good for
  • Agentic and tool-using workflows -- this is what it was built for
  • Coding and code review
  • Anything where you want reasoning quality without a 120B latency bill
Detail, limits and gotchas
Notes
  • Only 3B of the 35B parameters run per token, which is why it is the fastest capable model here.
  • It thinks before answering. Budget max_tokens accordingly -- see the reasoning-token warning below.
Deployment
  • 256 experts, 8 routed + 1 shared active per token
  • Limit is about 200 pages (approximately 131,000 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.

Upstream model card

Last check 12 min agoLatency 55msUptime 24h 100.0% (96 checks)

DeepSeek-R1 32B (distill)

Responding
reasoning

Reasoning-first. Shows its working.

ccs/deepseek-r1:32b
Size32B ArchDense Contextabout 45 pages (approximately 28,000 tokens)
thinks firstno tool calling
Good for
  • Maths, logic and step-by-step problems
  • Cases where seeing the reasoning is the point
Detail, limits and gotchas
Watch out
  • NO TOOL CALLING. Sending `tools=` returns HTTP 400. This is empirically verified, not a guess -- its Modelfile has no tools block. If you need tools plus reasoning, use Qwen3.6-35B-A3B-FP8.
  • This is NOT a search / embedding model, but it will still accept a 'turn this text into search numbers' request and hand back 5,120 numbers that look real and are not. A search tool built on them quietly returns poor results. Send embedding and semantic-search requests only to ccs/Qwen/Qwen3-Embedding-8B.
Notes
  • Writes long. It will use the whole max_tokens budget on reasoning if you let it.
Deployment
  • Limit is about 45 pages (approximately 28,000 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.

Upstream model card

Last check 12 min agoLatency 119msUptime 24h 100.0% (96 checks)

GPT-OSS 120B

Responding
reasoning

The largest model here. Reach for it when quality matters more than speed.

ccs/gpt-oss:120b
Size117B / 5.1B per token active ArchMoE Contextabout 190 pages (approximately 124,000 tokens)
thinks firsttools
Good for
  • Hard reasoning where you can afford the wait
  • Long-form writing and analysis
  • A second opinion when a smaller model looks wrong
Detail, limits and gotchas
Notes
  • Supports adjustable reasoning effort upstream; not currently exposed as a parameter here.
  • A large model that uses a full GPU on its own.
Deployment
  • Limit is about 190 pages (approximately 124,000 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.

Upstream model card

Last check 12 min agoLatency 207msUptime 24h 100.0% (96 checks)

Qwen3 32B

Responding
reasoning

Dense 32B with hybrid thinking.

ccs/qwen3:32b
Size32B ArchDense Contextabout 25 pages (approximately 15,500 tokens)
thinks firsttools
Good for
  • Reasoning tasks where you want a dense model rather than an MoE
  • Comparing against the MoE models on the same prompt
Detail, limits and gotchas
Notes
  • Thinks a LOT. Measured: 289 completion tokens to answer 'What is 2+2'. Give it room.
Deployment
  • Limit is about 25 pages (approximately 15,500 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.

Upstream model card

Last check 12 min agoLatency 109msUptime 24h 100.0% (96 checks)

Qwen3-VL 8B Instruct

Responding
vision

Send it images.

ccs/Qwen/Qwen3-VL-8B-Instruct-FP8
Size8B ArchDense vision-language Contextabout 100 pages (approximately 65,000 tokens)
toolsimages
Good for
  • Describing figures, charts, screenshots and scanned pages
  • Pulling text and structure out of an image
  • Any prompt that mixes an image with a question
Detail, limits and gotchas
Notes
  • Pass images as an `image_url` content part, the standard OpenAI vision shape. A base64 data URI works.
  • The only model here that accepts images. The others will ignore or reject them.
Deployment
  • Limit is about 100 pages (approximately 65,000 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.

Upstream model card

Last check 12 min agoLatency 57msUptime 24h 100.0% (96 checks)

Qwen3 Embedding 8B

Responding
embeddings

Turns text into vectors. It does NOT answer questions.

ccs/Qwen/Qwen3-Embedding-8B
Size8B ArchDense embedding model Contextabout 50 pages per item (approximately 33,000 tokens) Dimensions4096 Endpoint/v1/embeddings
Good for
  • RAG -- embed your documents, embed the question, retrieve the closest chunks
  • Semantic search over a corpus
  • Clustering, deduplication, similarity scoring
Detail, limits and gotchas
Watch out
  • Pass `encoding_format: "float"`. It used to be required -- without it the request failed with HTTP 400 -- and since the upstream upgrade on 2026-08-16 it succeeds either way. Keep sending it so your code does not depend on which upstream version happens to be deployed.
  • Sending a chat request to this model triggers an upstream cooldown: every caller of this model then gets HTTP 429 for roughly 30 seconds, not just the person who made the mistake. The model is not faulty when that happens and it recovers on its own. Measured 2026-08-16.
Not for
  • Chat or question answering. It has no answer to give -- it returns numbers, not text.
  • Do not send it to /v1/chat/completions. Use /v1/embeddings. A chat request to this model does not only fail for you -- it takes the model out for everyone for about 30 seconds.
Notes
  • Returns a 4096-dimension vector per input. Size your vector store for that.
  • Typical RAG shape: embed chunks once and store them, embed each query at ask time, retrieve top-k by cosine similarity, then pass those chunks to a chat model as context.
Deployment
  • Limit is about 50 pages per item (approximately 33,000 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.

Upstream model card

Last check 12 min agoLatency 75msUptime 24h 100.0% (96 checks)

Using the API

The gateway speaks the OpenAI API. Point any OpenAI-compatible client at the base URL below, authenticate with a Bearer token, and pass one of the model ids above as model.

Base URL https://llm.ccs.uky.edu/v1

This host issues no keys to people. For a durable personal key, use Locksmith. If you are running an Open OnDemand session or a batch job, a short-lived key is issued to the job automatically and revoked when the job ends. There is nothing to do.

curl
curl https://llm.ccs.uky.edu/v1/chat/completions \
  -H "Authorization: Bearer YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ccs/llama3.1:8b",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
Python (OpenAI SDK)
from openai import OpenAI

client = OpenAI(
    base_url="https://llm.ccs.uky.edu/v1",
    api_key="YOUR_KEY",
)

resp = client.chat.completions.create(
    model="ccs/llama3.1:8b",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)

Never share your key or embed it in published code. To check a key and compare models without installing anything, use Try it. See the full documentation for streaming, embeddings, and client integrations.