# National Compute: your first job

You are an agent running a researcher's first job on National Compute.
Read `https://nationalcompute.com/AGENTS.md` first: it covers the token,
`GET /api/whoami`, kubectl sign in, ssh keys, credits, how the market
prices capacity, and what a Marshall workspace remembers between
sessions (its Memory section). This page is the job itself. Every contract you
will touch is machine readable and needs no token to read:
`/api/vm/openapi.json`, `/api/k8s/openapi.json`,
`/api/market/openapi.json`, `/api/billing/openapi.json` — and the
starter recipes this page walks through are a catalog:
`GET https://access.nationalcompute.com/first-job/` lists every current
recipe as JSON (cluster kind, category, description, default file,
vendor variants).

## Before the job

1. **Token.** On your own machine your human saved it at
   `~/.config/nationalcompute/token`. `GET /api/whoami` with
   `Authorization: Bearer $(cat ~/.config/nationalcompute/token)`
   must return 200. Inside a Marshall workspace the daemon holds the
   credential; never ask for or mint a token there. The answer's
   `clusters` list names the cluster and its `kind`; the kind picks the
   section below. Its `org_kind` names the organization's kind. A
   `public_research` organization has an empty `clusters` list. Its
   members get whole GPU nodes from Marshall inside a workspace, so
   the cluster sections below do not apply to it: its section is
   `Public Research node`, after the Slurm one. Item 2 still does:
   a node request needs a balance for one full lease.
2. **Money.** `GET /api/billing/summary` returns `balance_ucredits`
   (1,000,000 = $1); `GET /api/billing/balance` returns it with
   `burn_ucredits_hr`, the hourly rate the meter last assessed for the
   organization. A Kubernetes job is refused at submit unless the
   balance covers two hours of it at your ceiling: with 8 GPUs and a $5
   ceiling that is $80. `purchases: true` means your human can top up
   on the console's Billing page. Tell them the number you need.
3. **Ceiling.** Ask your human for the most they will pay per GPU hour
   before you set a limit price or declare. Report the running rate back once nodes
   arrive: `GET /api/billing/balance` for the organization's metered
   rate (about ten minutes behind), `GET /api/k8s/market/spend` for a
   Kubernetes cluster's live market rate. For a
   VM cluster the market feed `GET /api/market/ticks` shows what
   winners are assessed, and your own charges land on
   `GET /api/billing/transactions`.

## Kubernetes cluster

**Connect.** You hold the token, so fetch the kubeconfig by token.
The portal resolves the name inside your organization, so another
organization's cluster can never answer:

    mkdir -p ~/.kube && curl -fsSL \
      -H "Authorization: Bearer $(cat ~/.config/nationalcompute/token)" \
      "https://nationalcompute.com/api/k8s/cluster/kubeconfig?cluster=<cluster>" \
      -o ~/.kube/<cluster>.yaml
    export KUBECONFIG=~/.kube/<cluster>.yaml

The setup script does the same and also installs kubectl and the sign
in plugin when missing. It reads the token from
`~/.config/nationalcompute/token` (or `NC_TOKEN`) and merges a
`<cluster>` context into `~/.kube/config`:

    curl -fsSL https://access.nationalcompute.com/setup.sh | sh -s -- <human's login email> <cluster>

Either way the sign in is a browser step only your human can do, and
nothing waits on you: run without a terminal, the script prints one
sign in link and exits while sign-in polls in the background (on the
manual path the first kubectl call prints the same link). Hand the
link to your human verbatim. They open it and approve once. When they
say they have, confirm with `kubectl --context <cluster> get nodes` —
it answers at once after approval; a fresh link means the code
expired, so relay that one. After that kubectl signs in silently while
you keep using it. An idle day needs one more approval — the next
kubectl call prints the link again, so run kubectl where you can read
its output. Name the context on every kubectl call.

**Know the cluster.** `GET /api/k8s/cluster` (the token's cluster, or
`?cluster=<name>`) returns `market` (true = the market grants nodes to
your jobs; false = a fixed cluster, skip the limit price), `gpus_per_node`,
`cpu_worker`, `shared_volume` (a volume at `/mnt/shared` on every
node), `gpu_vendor`, and `gpu_model` (the GPU class the cluster
trades). The GPU resource name in a manifest is `<gpu_vendor>.com/gpu`.

**Limit price** (market clusters only). First look at what the cluster
already holds: `kubectl get nodes` — an idle GPU node already present
(a base-load reservation, or an earlier grant still held) hosts the job
with no new bid, and reserved nodes serve requests before the market.
Only when the job needs more than what is held: `PUT /api/k8s/bid` with
`{"max_price_per_gpu_hour": <ceiling>}` sets the most you will pay per
GPU hour. Read what wins off the market's own numbers first — the
price to win ladder (`GET /api/k8s/market/ladder`) and the ticks
(`GET /api/market/ticks`); a ceiling too low to ever clear the market
is refused (`422 bid-too-low`).

**Manifests.** The sample files live at
`https://access.nationalcompute.com/first-job/`, and a GET on that URL
itself returns the current catalog as JSON — every recipe with its
description, default file and vendor variants. Where a manifest has a
vendor variant, try `<name>-<gpu_vendor>.yaml` first; a 404 means the
default `<name>.yaml` is the one to apply. Pass a file's URL straight
to `kubectl apply -f`.

**Which recipe.** When the hand off prompt names a recipe, that is the
answer. Do not ask when the answer is already in the prompt. If your
human did not say which, run Train: one node, nothing to build, done
in minutes. Run Serve only when they asked for a model endpoint. Run
Something else only when they gave you a container image and a
command.

### Train

Trains a small GPT with PyTorch DDP on all GPUs of one node, in the
stock PyTorch container. Nothing to build.

    kubectl --context <cluster> apply -f https://access.nationalcompute.com/first-job/nanogpt-job-<gpu_vendor>.yaml
    kubectl --context <cluster> get pods -w
    kubectl --context <cluster> logs -f job/nanogpt

If the vendor variant is a 404, apply the default,
`https://access.nationalcompute.com/first-job/nanogpt-job.yaml`.

On a market cluster the job waits suspended until its node is granted;
the first node takes 10 to 15 minutes to build and join. `kubectl get
pods` stays empty while the job waits, which is normal. The pod shows
`ContainerCreating` while the image pulls, then `Running`; the log
prints lines like `iter 100: loss 2.4712, time 380.12ms`. Stop with
`kubectl --context <cluster> delete job nanogpt`. Reclaims and their
verdicts are in AGENTS.md's operating loop.

### Serve

Two serving recipes, each a large open model as an OpenAI-compatible
endpoint on one GPU node. Each is offered only where its variant exists
for the cluster's vendor: `<recipe>-<gpu_vendor>.yaml` 404s elsewhere.
Run one serving recipe at a time. The Getting Started Serve tab and a
hand off that says "run its Serve recipe" mean Kimi K3; Marshall's
Serve card names `inkling-serve`.

#### Kimi K3

Behind a router, one engine replica on one GPU node.

1. With a shared volume, stage the 1.5 TB weights first. The staging
   Job runs on the CPU worker while the GPU node is still arriving. The
   volume needs about 1.5 TB free:

        kubectl --context <cluster> apply -f https://access.nationalcompute.com/first-job/kimi-k3-weights-to-shared.yaml
        kubectl --context <cluster> logs -f job/kimi-k3-weights

2. Launch the router and the worker:

        kubectl --context <cluster> apply -f https://access.nationalcompute.com/first-job/kimi-k3-serve-<gpu_vendor>.yaml
        kubectl --context <cluster> get pods -w

   The router pod reads `Running 1/1` within a minute. The worker reads
   `SchedulingGated` while the market grants its node and `Pending`
   while the node joins, then `Init:0/2` while the weights stage,
   copy from the volume, or download, then `Running` with `READY 0/1`
   while the engine loads (about 5 minutes, or 10 on the cluster's
   first start). Without staged weights the worker downloads 1.5 TB
   onto the node, which takes 20 minutes to a few hours.

3. The model is up at `READY 1/1`. Send a request:

        kubectl --context <cluster> port-forward svc/infera 8000:8000 &
        curl -s localhost:8000/v1/chat/completions -H 'content-type: application/json' \
          -d '{"model":"moonshotai/Kimi-K3","messages":[{"role":"user","content":"hi"}]}'

4. Optional load test, about 17 minutes, from inside the cluster:

        kubectl --context <cluster> apply -f https://access.nationalcompute.com/first-job/kimi-k3-loadtest.yaml
        kubectl --context <cluster> logs -f job/kimi-k3-loadtest

   Do not send your own requests during the run. The results also
   appear on the Job's page under Workloads in the console.

#### Inkling

Thinking Machines' open-weights model: 975B parameters, 41B active,
up to 1M tokens of context, reasoning and tool calling, served by the
stock inference engine with no router. The checkpoint is about 530 GiB,
a third of Kimi K3's, so a 1 TiB shared volume holds it.

1. With a shared volume, stage the weights first. The staging Job runs
   on the CPU worker while the GPU node is still arriving. The checkpoint
   differs by GPU vendor, so the Job has a vendor variant like the engine:

        kubectl --context <cluster> apply -f https://access.nationalcompute.com/first-job/inkling-weights-to-shared-<gpu_vendor>.yaml
        kubectl --context <cluster> logs -f job/inkling-weights

   If the vendor variant is a 404, apply the default,
   `https://access.nationalcompute.com/first-job/inkling-weights-to-shared.yaml`.

2. Launch the engine:

        kubectl --context <cluster> apply -f https://access.nationalcompute.com/first-job/inkling-serve-<gpu_vendor>.yaml
        kubectl --context <cluster> get pods -w

   The `inkling` pod reads `SchedulingGated` while the market grants
   its node and `Pending` while the node joins, then `Init:0/1` while
   the weights copy from the volume (about a minute) or download from
   the Hub (10 to 20 minutes), then `Running` with `READY 0/1` while
   the engine loads and builds its kernels: budget 15 minutes on a
   cluster's first start, less once its compile caches on the shared
   volume are warm. `kubectl --context <cluster> logs -f deploy/inkling -c main`
   follows the engine.

3. The model is up at `READY 1/1`. It reasons before it answers, so
   give a request room:

        kubectl --context <cluster> port-forward svc/inkling 8000:8000 &
        curl -s localhost:8000/v1/chat/completions -H 'content-type: application/json' \
          -d '{"model":"thinkingmachines/Inkling","messages":[{"role":"user","content":"hi"}],"max_tokens":1024}'

   The answer is the message's `content`; the thinking arrives in the
   message's separate reasoning field. Tool calling works the OpenAI
   way (`tools` in the request).

Stop either with `kubectl --context <cluster> delete -f <the serve manifest URL>`.
Staged weights stay on the shared volume for the next start.

### Post-train

Fine-tune an open chat model into a financial risk analyst on a folder
of documents, then talk to the result — end to end on one node, one
limit price. The starter's corpus is fully synthetic (generated filings with
expert risk analyses, Apache-2.0), so the demo never touches real
books; the same recipe shape takes any folder of documents. Needs a
cluster with a shared volume.

1. Tune. The Job stages the document folder to
   `/mnt/shared/finance-analyst/docs` (open it — that folder is the
   training set), LoRA-tunes the base model on every GPU of the node,
   and writes the merged model to `/mnt/shared/finance-analyst/model`.
   The log ends with the tuned model's memos for two held-out test
   documents printed next to their reference analyses — read them
   before serving; a defective tune shows up there first:

        kubectl --context <cluster> apply -f https://access.nationalcompute.com/first-job/finance-analyst-tune-<gpu_vendor>.yaml
        kubectl --context <cluster> logs -f job/finance-analyst-tune

   If the vendor variant is a 404, apply the default,
   `https://access.nationalcompute.com/first-job/finance-analyst-tune.yaml`.

2. Serve. Safe to apply at the same time as the tune Job: a launcher on
   the CPU worker holds the engine at zero replicas until the tune has
   succeeded, then scales it up and the engine takes the node the
   trainer just freed — one node, serially, the same limit price funds
   both phases. While it waits, nothing holds (or bills) a GPU node:

        kubectl --context <cluster> apply -f https://access.nationalcompute.com/first-job/finance-analyst-serve-<gpu_vendor>.yaml
        kubectl --context <cluster> get pods -w

3. Talk to it at `READY 1/1` — a private, cluster-internal endpoint:

        kubectl --context <cluster> port-forward svc/finance-analyst 8000:8000 &
        curl -s localhost:8000/v1/chat/completions -H 'content-type: application/json' \
          -d '{"model":"finance-analyst","max_tokens":700,
               "chat_template_kwargs":{"enable_thinking":false},
               "messages":[
                {"role":"system","content":"You are the organization'"'"'s financial risk analyst. Given a filing or financial document, write a concise risk memo: overall severity, risk categories, a grounded analysis, financial impact, key metrics, and critical dates. Report only what the document supports — omit anything it does not state."},
                {"role":"user","content":"<paste a document from /mnt/shared/finance-analyst/docs>"}]}'

   Query it exactly like that: the system prompt is the one the model
   was tuned with, `enable_thinking:false` matches how it was tuned
   (the base is a hybrid-thinking model), and sampling is left unset on
   purpose — the tuned model ships its own decoding defaults. Never
   force `temperature` to 0: greedy decoding makes this model repeat
   itself.

Rerunning the tune: a serve pod keeps the weights it loaded at startup,
so a pod that started before the new tune finished is serving the
PREVIOUS model — delete the Deployment before re-tuning and re-apply it
after (or `kubectl --context <cluster> rollout restart
deploy/finance-analyst` once the new model lands). The tune grades
itself on held-out documents at the end of its log (also kept on the
shared volume at `/mnt/shared/finance-analyst/last-tune.log`) and
refuses to bless a defective model: the Job fails instead of writing
the `.ready` marker, and the serve leg keeps waiting.

Stop with `kubectl --context <cluster> delete -f <the serve manifest
URL>` and `kubectl --context <cluster> delete job finance-analyst-tune`.
The tuned model and the document folder stay on the shared volume.

### Something else

Your human's own container as one Job on a whole node. Set the image
and the command; the pod asks for every GPU of one node:

    apiVersion: batch/v1
    kind: Job
    metadata:
      name: my-job
    spec:
      backoffLimit: 0
      template:
        spec:
          restartPolicy: Never
          terminationGracePeriodSeconds: 60
          tolerations:
            - key: <gpu_vendor>.com/gpu
              operator: Exists
              effect: NoSchedule
          containers:
            - name: main
              image: YOUR_IMAGE
              command: ["sh", "-c", "YOUR_COMMAND"]
              resources:
                limits:
                  <gpu_vendor>.com/gpu: <gpus_per_node>
              volumeMounts:
                - name: shm
                  mountPath: /dev/shm
          volumes:
            - name: shm
              emptyDir: { medium: Memory, sizeLimit: 8Gi }

Every pod must request a multiple of `gpus_per_node` GPUs (a partial
node is refused at submit). The exception is a cluster with reservation
packing enabled. There a pod may request fewer GPUs than one node. Such
a pod runs only on your reservation and never bids. For work that must
survive a reclaim,
checkpoint to `/mnt/shared` and exempt reclaims from the Job's retry
budget:

    spec:
      podReplacementPolicy: Failed
      podFailurePolicy:
        rules:
          - action: Ignore
            onPodConditions:
              - type: DisruptionTarget
                status: "True"
          - action: Ignore
            onExitCodes:
              operator: In
              values: [143]

`podFailurePolicy` needs `restartPolicy: Never`, which the skeleton
above sets. AGENTS.md's operating loop explains the reclaim itself.

## VM cluster

No kubectl. Nodes arrive as ssh hosts.

1. Register your human's public key with `POST /api/vm/orgs/{org_id}/ssh-keys`
   (`org_id` from whoami) if `GET .../ssh-keys` does not list its
   fingerprint. Keys install on future grants only.
2. Declare: `PUT /api/vm/capacity` with `{"max_gpus": <gpus_per_node>,
   "max_price_per_gpu_hour": <ceiling>}`, echoing the read's `version`
   as `expected_version`.
3. Poll `GET /api/vm/capacity` until an allocation reads `active`; its
   `public_ip` and the payload's `environment.ssh_user` are the login:
   `ssh <ssh_user>@<public_ip>`. Run the workload there in a container
   of your human's choice.
4. Done: `PUT /api/vm/capacity` with `max_gpus: 0` and the same
   `max_price_per_gpu_hour` (the PUT needs both fields). Billing stops
   with the release. A released node is destroyed with its local disk.

## Slurm cluster

The sign in is your human's: they run the setup script (as above, with
the cluster name) and then `slurm-login`, which refreshes an ssh
certificate valid for 24 hours. After that `ssh <cluster>` from the
same machine lands you on the login node. There:

    curl -fsSL https://access.nationalcompute.com/first-job/prepare.sh | sh
    sbatch --nodes=1 ~/first-job/nanogpt.sbatch
    squeue --me
    tail -F nanogpt-<jobid>.out

## Public Research node

No cluster and no kubectl: one whole node, requested from Marshall
inside a workspace and reached over ssh. `GET /api/whoami` reads
`org_kind: public_research` and an empty `clusters` list.

1. **Request.** Only Marshall's workspace tools do this
   (`public_research_request`; there is no API route and no MCP tool).
   The request ends in a confirmation card your human approves, and it
   needs a balance for one full lease (item 2 above). Then wait: the
   pool watch reports the queue position and the estimated wait (an
   estimate that assumes every lease ahead runs its full length) and
   wakes the session when the node is assigned. Never poll.
2. **Connect.** The status (`public_research_status`) names the node's
   `public_ip` and `ssh_user`. The workspace key is already on the node
   and its host key is pinned, so `ssh <ssh_user>@<public_ip>` works as
   is from a shell pane or `shell_run`. Port 22 is the only way in.
3. **Train.** One line on the node starts the same nanoGPT run the
   clusters get, in the stock PyTorch container for the node's GPUs,
   detached from the ssh session:

        ssh <ssh_user>@<public_ip> 'curl -fsSL https://access.nationalcompute.com/first-job/nanogpt-node.sh | sh'
        ssh <ssh_user>@<public_ip> 'tail -F ~/first-job/nanogpt.out'

   The script fetches the project under `~/first-job` on the node and
   runs it with `docker`; when docker is missing from the node it says
   so and stops. The first start pulls the image (several minutes).
   Then the log prints lines like `iter 100: loss 2.4712, time
   380.12ms`. The run is time-boxed to about an hour and ends on its
   own. Stop it early with `docker rm -f nanogpt` on the node.
4. **Keep the results.** The node's disk is scratch and dies with the
   lease. Copy the checkpoint and the log into the workspace before
   the release:

        scp -r <ssh_user>@<public_ip>:first-job/nanoGPT/out-shakespeare-char .
        scp <ssh_user>@<public_ip>:first-job/nanogpt.out .

5. **Release.** The meter runs until the node is released, and the run
   ending does not release it. Never release on your own: say the run
   is done and the meter is still running, then ask your human whether
   to release (`public_research_release`, a confirmation card).

Serve, Post-train and Something else are cluster recipes (Kubernetes
manifests). On a node, run your human's own container the same way
with `docker run` on every GPU of it. Nothing here needs a shared
volume and nothing spans two nodes.

## Report back

Tell your human: the GPU class the job ran on (`gpu_model`), when the
node arrived, the first loss lines or the first completion, the assessed
rate per GPU hour, and how to stop it.
Leave the cluster the way they want it: a standing limit price keeps
pulling nodes for any GPU job they apply later; `PUT /api/k8s/bid` with `0`
withdraws it.
