No key yet. The examples below show YOUR_KEY until you paste one.

Where your key goes, and what we cannot protect you from

What this page does. Your key stays in this browser tab, in the field above and nowhere else. When you press Send, the browser calls the endpoint you chose (/v1/chat/completions, /v1/embeddings or /v1/messages) on this host directly with an Authorization: Bearer header. An image you choose for the vision test is read by the browser and travels only inside that request body; it is never uploaded to or saved on this host, and the page forgets it when you remove it, switch mode, clear the key or reload. In the tool-calling test the tools (a clock and a calculator) run in your browser too. The key is never sent to this web page as form data, never put in a URL, never written to a cookie, and never stored in localStorage or sessionStorage. Leaving or reloading the page discards it; Clear removes it immediately.

What the gateway records. Request metadata only: time, model, status, token counts, latency, client address, and the last four characters of the key so an operator can trace one key during an incident. The full key is never written to a log.

What we cannot protect you from. A browser extension can read anything on this page, including the key field. So can anyone watching your screen, in person or over a shared screen or a recording. Copy buttons put the key on your system clipboard, where other applications can read it. A snippet you paste into a terminal lands in your shell history (~/.bash_history) unless you prefix the line with a space or use the environment-variable form. If any of that is a concern for the key you hold, use the terminal examples with the key in an environment variable rather than pasting it here.

If a key is exposed, revoke it at https://locksmith.ccs.uky.edu and generate a new one. Keys are cheap; a leaked one is not.

Try Ctrl+Enter sends

Reasoning models spend tokens thinking before they answer

A reasoning model uses part of your max_tokens budget on internal reasoning. If the budget is small it can spend ALL of it thinking and return an empty string -- which looks like a broken service but is not. Measured here: ccs/qwen3:32b used 289 completion tokens to answer "What is 2+2". At max_tokens=24 the same question returned "". Ask reasoning models for at least 2,000 tokens, or use a non-thinking model such as ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink or ccs/llama3.1:8b.

Run the same thing from a terminal

These update as you change the test, key, model and prompt above.
curl: any machine with a shell
curl https://llm.ccs.uky.edu/v1/chat/completions \
  -H "Authorization: Bearer YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink",
  "messages": [
    {
      "role": "user",
      "content": "What is 2+2? Answer in one short sentence."
    }
  ],
  "max_tokens": 512
}'

A copied line goes into your shell history. Prefix it with a space, or use the Environment tab, if that matters to you. Full base URL: https://llm.ccs.uky.edu/v1 · no key yet? Get one from Locksmith · more clients in the documentation · what each model is for, on the status page.