# National Compute — agent orientation

You are talking to a market, not a reservation system. Nothing here
books capacity outright: you state what you want and the most you'll
pay, and the market decides what you hold, tick by tick.

## Where the spec is

- `GET /api/vm/openapi.json` — the VM capacity contract (OpenAPI 3.1):
  every endpoint, field, rule, and error code. No auth needed.
- `GET /api/market/openapi.json` — the market feed contract (the
  clearing-price endpoints both kinds read; OpenAPI 3.1). No auth
  needed.
- `GET /api/k8s/openapi.json` — the limit price contract for Kubernetes
  clusters. No auth needed.
- `GET /api/billing/openapi.json`: the billing contract. Your
  organization's ledger (charges per node, daily totals, the balance
  with the hourly burn rate and its history, unit prices), the hosted
  checkout, card setup and payment portal pages you
  hand to your human, a spend limit you may only tighten, and the auto
  reload setting (a standing card charge; your human approves any
  change to it before you write it). No auth needed.
- `GET /api/org/openapi.json`: the organization contract. Your
  organization's roster, join rules, outstanding invitations, the
  metadata of every API token it holds, each token's activity trail,
  and a token's own revoke. No auth needed for the contract.

Read the spec before writing; it is the authority whenever this page
and the spec disagree.

**Or skip the REST plumbing entirely**: the platform hosts an MCP
(Model Context Protocol) server at `https://nationalcompute.com/mcp` —
the same surface as tools, plus server-side Kubernetes access (no
kubeconfig, no kubectl on your machine). Auth is an OAuth browser
sign-in by your human (no API token to collect), every action is
attributed to them, and every spending or mutating tool is two-phase:
the first call returns a preview and a single-use `confirm_token`
(300 s); the same call with the token executes. Show your human the
preview in between. Inside a Marshall workspace (see Memory below) the
console draws a `cluster_apply` preview as a launch card titled
`Confirm Launch - <Kind>`, with the YAML behind a fold and the buttons
Launch and Cancel. Your human answers on the card. The tool result you
receive is unchanged. A preview that also carries `human_confirmation`
(`{required: true, reason}`) spends money or destroys data: today
`vm_swap`, a vm `bid_write` that would release granted nodes,
`cluster_delete` of a PersistentVolumeClaim, Namespace, VolumeSnapshot
or VolumeSnapshotContent, `cluster_delete` of a StatefulSet whose
volume claims delete with it, and a `cluster_apply` that makes a
StatefulSet delete its claims. A PersistentVolume delete is refused
today as an unknown kind and
joins the class if that kind is ever added. Relay the reason and wait
for your human's own answer to that call. No standing, remembered or
full access approval covers it. Inside a Marshall workspace a new
session starts with Full Access on. Shell commands, pane typing,
ordinary grid actions and shared memory saves then run without an
approval card until an owner turns Full Access off. Workspace restarts,
signatures, the Public Research node request card and this human
confirmation class still ask. A workspace
member token cannot run such a call at all
(`403 workspace-human-class-refused`). Full page:
`https://docs.nationalcompute.com/api/mcp/`. Everything below remains
the REST path's ground truth, and the MCP tools follow the same market
rules. The reads documented below reach the MCP through `grid_api
{action: get, path, params}` by their REST paths: the Workloads page,
Run History, a workload's cost, the job page and its cost and network
panels, the Base Load book, the kubeconfig, the storage volumes, the
token list and both activity trails, beside the market, billing and
organization feeds. A `{jobid}` or `{token_id}` segment takes the REST
spelling or the templated key with the value in `params`. The bodies
and the refusals are the REST routes' own.

Running the human's **first job** — connect, price, launch, watch, for
a Kubernetes, VM or Slurm cluster — is `/FIRST-JOB.md`, the page the
console's Getting Started hand off points at. This page is the ground
rules it assumes. Its starter recipes are enumerable, no auth:
`GET https://access.nationalcompute.com/first-job/` returns the current
catalog as JSON (each recipe's cluster kind, category, description,
default file and vendor variants).

## What to collect from your human first

Only two things gate everything below, and only a person can provide
them:

1. **An org API token**, for an agent on its own machine. Mint it in a
   browser on the portal's Getting Started page. One token acts for its
   organization on every customer API: reading and declaring capacity,
   setting a limit price, keys, the price feed, and the billing records.
   There is no API to mint one. Ask your human to do it and to save it
   where you can read it without the secret ever entering chat or shell
   history, e.g.:

        mkdir -p ~/.config/nationalcompute && printf 'token: ' && \
        read -rs t && printf '%s' "$t" \
          > ~/.config/nationalcompute/token && \
        chmod 600 ~/.config/nationalcompute/token && unset t

   then read it back per call with
   `$(cat ~/.config/nationalcompute/token)`.
2. **Spending authority** — a limit price is real money on prepaid
   credits, metered every hour you hold nodes. Get an explicit maximum
   from your human before your first declare or limit price, report the
   running rate back, and treat every later ceiling change as a new
   spending decision, not a tuning knob you own.

**Discover the rest — don't ask.** With the token in hand,
`GET /api/whoami` answers the remaining bootstrap questions in one
authenticated read: the org you act as (`org_id`, its display name
`org_display` and its kind `org_kind`), your scopes, whether the token
is bound to one cluster, and every cluster it can address as `{name,
kind, gpu_vendor, gpu_model, market, gpus_per_node, shared_volume,
cpu_worker}`. `org_kind` reads `standard` for an ordinary organization.
`public_research` names a per user organization created at first sign
in for a `.edu`, `.mil` or `.gov` address that matches no join rule and
belongs to no organization yet (see the organization paragraph below). The MCP tool `grid_whoami` carries the same field.
`gpu_model` is the GPU class the cluster's site sells (null while
the registry has none). It tells you what you price before you set
your first limit price or read capacity. `gpu_vendor` is its maker.
The manifest resource name is `<gpu_vendor>.com/gpu`. `market` true
means the cluster trades on the market: a limit price or a capacity
declaration applies to it. `market` false is a fixed cluster the
platform sizes; it has no limit price to set. `market` null means the
posture is not known yet (the cluster record has not landed); read
again later. `shared_volume` says a
shared filesystem is bound to the cluster (null when the storage
record could not be read) and `cpu_worker` says it carries a CPU only
node. Kind tells you which surface to drive (each surface refuses the
other kind's token). Go back to your human only when several clusters
of the right kind exist and context (a kubeconfig context name, the
task itself) doesn't already pick one — that is an intent question,
not a discovery question. For a Kubernetes cluster,
`GET /api/k8s/cluster` then gives the facts the console shows: nodes
and GPUs held, `gpus_per_node`, `gpu_model`, whether it trades on the
market, the CPU worker and shared volume, the injected NCCL
environment, and the kubeconfig URL and endpoints. For a VM cluster the
capacity read carries `gpus_per_node`, `gpu_model`, your allocations,
and `environment` (the ssh user and the fabric settings).
`GET /api/k8s/nodes` is the Cluster Overview board: every node with
`verdict`, the platform's own health word (`""` healthy, `unobserved`,
`unreachable`, `NotReady`, `Node Fix In Progress`, `unhealthy`,
`tainted`, `cordoned`, `missing`, or a lifecycle status), and
`GET /api/k8s/nodes/{node}` is one machine. `verdict.observed` is
false when the scheduler feed did not answer on that read; a node
nothing judged reads `unobserved`, never healthy. Read `island_stale`
and `feed_error` on `GET /api/k8s/cluster` before acting on a red
verdict. A dark site makes every verdict stale. On the MCP the same
two bodies are `grid_api get /api/k8s/nodes` and
`grid_api get /api/k8s/nodes/{node}`.
The organization itself is two reads: `GET /api/org/members` returns
the roster, the auto join rules, and the outstanding invitations.
`GET /api/org/tokens` returns every token as metadata (name, scopes,
cluster binding, expiry, creator). The secret is never on the wire.
Your own trail is `GET /api/org/tokens/self/activity`: every write you
made with this token (declares, swaps, limit prices, key changes) and the
token's own mint and revoke rows, newest first. Writes are always
recorded. The capacity GET and the ssh key list are recorded on every
call too; those read rows stay out of the trail unless you pass
`reads=1`. Other reads are not recorded. Read the trail back to your
human when a task ends, so they see what changed in their name. When
your human asks you to end your own access, or you believe the token
leaked, `DELETE /api/org/tokens/self` revokes it at once; the next call
with it is 401. Inviting a member, adding a rule, minting a token, and
revoking any other token stay console actions for your human. So does
buying Base Load (a reservation). Destroying the shared NFS volume (the
Storage page's volume destroy) is console only. A Kubernetes
PersistentVolumeClaim delete is allowed through `cluster_delete` under
the human confirmation class. No agent path mints a token; the human
does it on the console's `/org/<org>/api-keys` page.

A `public_research` organization is created with no clusters and a
welcome credit airdrop as its first ledger row (a deposit of source
`airdrop` on the billing reads; the platform sets the amount). Its
`clusters` list is empty, so none of the cluster surfaces below apply
to it. Its members buy further credits on the console Billing page
like any organization. A node request needs a
balance for one full lease (the quote names the rate and the lease
length); until then the request is refused with `balance_too_low` and
the `required_ucr` figure. Its members hold one whole GPU node at a
time, requested from Marshall inside a Marshall workspace and used over
ssh. The
request, status and release tools exist only in that workspace. The
routes behind them admit only the workspace daemon's own credential.
They have no MCP twin. Marshall never releases a node on its own. The
member asks for the release. The first job on a node is FIRST-JOB.md's
`Public Research node` section: the catalog's `nanogpt-node` recipe, a
script run on the node over ssh. A join rule for a whole `.edu`, `.mil` or
`.gov` institution (`stanford.edu`) is refused. A department or lab
domain (`lab.cs.stanford.edu`) or an individual address is accepted.
Full page: `https://docs.nationalcompute.com/public-research/`.

For a Slurm cluster the MCP server is the read side and the login node
is the write side. `grid_whoami` names the cluster with its login
endpoint (`ssh_host`, `ssh_port`, `ssh_user`, `ssh_auth`; `null` while
unresolved). `slurm_read` answers what the console's status and job
pages show: node states and GPU allocation, the live queue with each
pending job's reason, GPU hours this month per user, the shared `/home`
usage, one job's record and queue position, its log tail or a grep over
it. `metrics_read` reads a Slurm node's telemetry by the node name
`slurm_read status` uses. Submitting, cancelling and everything else
Slurm runs over ssh on the login node, in your own shell. The status,
job, queue and log reads have no token route; the MCP is their only
agent surface. The login facts ride the token route too:
`GET /api/whoami` carries them on every `slurm` entry. Full page:
`https://docs.nationalcompute.com/api/slurm/`.

Also know your workload's GPU shape — how many GPUs each pod or node
will request — before pricing anything: whether a bid can clear
depends on that shape at least as much as on the number (see below).

## Before you start

Walk this checklist before you read prices or touch capacity. Resolve
each miss with your human. Every item is cheap to check now and a
confusing failure in the middle of a run.

1. **Bearer token on disk.** Look for the token at the expected path,
   `~/.config/nationalcompute/token`. If it is missing, ask your human
   to mint one on the portal's Getting Started page and hand them the
   `read -rs` command above to run — the secret goes straight into
   the file without you, the chat, or shell history ever seeing it.
   Then confirm it works: `GET /api/whoami` must return 200.
2. **Kubernetes tooling** (whenever `whoami` lists a cluster of kind
   `k8s`): on your own machine check that `kubectl` is installed and
   that kubelogin is reachable as its `oidc-login` plugin with
   `kubectl oidc-login --version`.
3. **Kubernetes login** (same condition): run the setup script — it
   installs both tools above when missing, merges a `<cluster>` context
   into `~/.kube/config` without touching other contexts, and is
   idempotent, so run it even when the tools check passes:

        curl -fsSL https://access.nationalcompute.com/setup.sh \
          | sh -s -- you@example.com <cluster>

   Fill in their login email and the cluster name from `whoami`. The
   script reads your token from `~/.config/nationalcompute/token` (or
   `NC_TOKEN`) and fetches the kubeconfig by token from
   `GET /api/k8s/cluster/kubeconfig?cluster=<cluster>`, the URL
   `GET /api/k8s/cluster` returns as `kubeconfig_url`. That route
   resolves the name inside your organization, so another organization's
   cluster can never answer. The public onboarding path refuses a name
   that two sites carry with `409 ambiguous-cluster`. The login is a
   browser step only your human can do, and the script never waits on
   it: run without a terminal, it prints one sign in link and EXITS
   while sign-in keeps polling in the background. Read the script's
   full output and hand the link to your human verbatim. Once they say
   they approved, confirm with `kubectl --context <cluster> get nodes`
   — it answers at once after approval; a fresh link means the code
   expired, so relay that one the same way. Never run the script or
   kubectl somewhere its output is hidden from you: a command that
   seems hung is showing a sign in link your human has not seen. The
   script and everything after it are yours to run — only the link
   crosses to your human.
4. **An ssh key registered** (whenever `whoami` lists a cluster of
   kind `vm`): granted nodes are reachable only through the org's
   registered ssh keys, and raising `max_gpus` with none registered is
   refused (`422 no-ssh-keys`). `GET /api/vm/orgs/{org_id}/ssh-keys`
   (`org_id` from `whoami`) lists fingerprints; compare with your
   human's local key (`ssh-keygen -lf ~/.ssh/id_ed25519.pub`). If
   theirs is absent, add it — with their OK — via
   `POST /api/vm/orgs/{org_id}/ssh-keys`, body
   `{name, public_key[, cluster]}`. Keys install on FUTURE grants
   only, so register before declaring capacity, not after.
5. **Credits loaded.** Capacity bills against prepaid credits:
   `GET /api/billing/summary` returns the balance (`balance_ucredits`,
   1,000,000 = $1) — a zero or negative balance means your human loads
   credits before anything fills. You can start that for them:
   `POST /api/billing/checkout` with `{usd}` returns a hosted payment
   page URL. Hand the URL to your human and wait. The page never
   offers a saved card: your human enters the card, so the URL alone
   cannot pay. The deposit lands when they pay, so read the summary
   again afterwards. Never enter card details yourself, and never pick
   the amount alone (it is a spending decision, see 2). `POST /api/billing/setup` returns a
   hosted page that saves a card, and `POST /api/billing/portal` opens
   the payment portal for receipts and card changes. Card payments are
   capped at `card_max_usd` total per rolling `card_window_days` days
   per organization; the summary's `card_window_remaining_usd` is what
   a card can still carry right now and `card_window_blocked` is true
   when nothing can. Larger amounts go by wire transfer. A non-null
   `billing_hold` on a VM capacity read means billing needs resolving
   before anything will fill. The monthly spend limit
   (`spend_limit_ucredits` on the summary) is your human's cap on you:
   `PUT /api/billing/limit` with `{usd}` sets one where none exists or
   lowers it. Raising or clearing it is refused (`403
   session-required`) and stays a console action for your human.
   `spend_limit_used_pct` on the summary is the month's charges as a
   share of the limit (the console warns at 75 and turns critical at
   99). Auto reload, a standing charge of `amount_usd` on the saved card
   whenever the balance falls under `threshold_usd`, is written by
   `PUT /api/billing/reload` with `{enabled, threshold_usd, amount_usd}`
   under the same posture as the limit: a token may disable it, lower
   the amount or raise the threshold. Enabling it, raising the amount
   or lowering the threshold is refused (`403 session-required`) and
   stays your human's console action, or the MCP `billing_write` tool
   where your human approves the preview. It is a spending
   authorization that outlives the conversation: obtain your human's
   explicit approval before changing it in any direction. A write that
   changes nothing is a no op and never re arms a paused reload. The
   summary's `auto_reload.last` reports the newest attempt
   whether or not it is enabled now, and `auto_reload.cooldown_until`
   says until when a failed charge pauses new ones (saving the setting
   again re arms it). When the card cap refuses an amount, the summary's
   `wire` object (present while `purchases` is true) carries the wire
   transfer rail: bank, beneficiary, your organization's `reference`
   for the current UTC day (it MUST ride the wire's memo field; read it
   on the day of the wire) and the address to notify. On the MCP the
   same five actions are the `billing_write` tool: `checkout`, `setup`
   and `portal` answer `{url}` in one call; `limit` and `reload` are two
   phase and carry `human_confirmation` on the preview, which your
   client shows your human and never auto approves. Read
   `spend_limit_enforced` next to it. When it is false the platform
   records the limit and does not act on it for this organization's
   clusters, so the cap is your human's instruction to you and nothing
   else stops you. `billing_hold_enforced` says the same about a
   balance at or below zero. On Kubernetes the balance also gates
   launches: unless it covers **two hours** of the workload at your
   ceiling (2 × ceiling × the job's GPU total), the submit is refused
   outright — admission denial `BalanceTooLow`, naming the exact
   shortfall. The market re-runs the same check at every tick while a
   request waits; a shortfall there only leaves the request pending
   under the same verdict, and a top-up clears it at the next tick.
   Replacement pods of an admitted workload are exempt at apply, so a
   balance dip never blocks a running job's pod recreation. Once nodes
   arrive, `GET /api/billing/balance` returns the balance with
   `burn_ucredits_hr`, the hourly rate the meter last assessed across
   the organization (about ten minutes behind; `as_of` names the
   metering window), folded by kind, cluster and unit price, and
   `GET /api/billing/balance/history` the balance over time. Report
   those to your human; never estimate spend as a node count times a
   price read elsewhere. Where the organization can pay by card it can
   also issue itself an invoice for credits, paid by wire through its
   accounts payable team. `GET /api/billing/invoices` lists the
   organization's invoices and its billing profile.
   `GET /api/billing/invoices/{id}.pdf` is the document. On the MCP
   `billing_read view=invoices` reads the same. `invoice_write` has
   three actions. `action=profile` sets the billing profile (name,
   address lines, optional email) in one call. `action=issue` is two
   phase: a preview, then the same call with `confirm_token`.
   `action=cancel id=<id>` withdraws one open invoice in one call. The
   invoice id goes in the wire's memo field. The credits land when the
   wire does. Issuing moves none. An organization holds at most ten open
   invoices. Past that the issue is refused with `409 open-limit`. A
   cancel of a paid, void or already cancelled invoice is refused with
   `409 not-open`. Cancel only an invoice your human names. The platform
   skill `nc-request-invoice` carries the exact flow and is installed
   for every member by default; read it with `skill_read` before the
   first invoice. Where the feature is absent the invoice routes read
   `404`. Marshall model usage is charged to the organization's credits
   at cost; the billing summary's `inference` block carries the charged
   amounts and any free allowance that applies (`free_tier_usd`, 0 =
   charged from the first call). Full page:
   `https://docs.nationalcompute.com/billing/`.
6. **Capacity already held.** List the cluster's nodes and read
   `GET /api/k8s/market/demand` — its `reserved` band is capacity your
   organization pays for OUTSIDE the market (a base-load reservation;
   the workloads read's market block carries the same reserved count).
   A workload that fits onto idle held or reserved nodes needs NO limit
   price: requests seat on reserved nodes before the market, and only
   the remainder becomes market demand. Set (or keep) a nonzero ceiling
   only when your human wants nodes BEYOND what they already hold — the
   ceiling is a standing order for that overflow, so name what it funds
   when you propose it. On a cluster with reservation packing enabled a
   job that asks for fewer GPUs than one node is packed onto a reserved
   node beside your other small jobs. It never bids on the market. While
   the reserved GPUs are full it waits with the reason `waiting for
   reserved GPUs`. No limit price changes that wait.

   Asking the platform team for what self serve cannot do (a Base Load
   reservation, a capacity watch, an unsupported ask, feedback) is the
   `escalate` MCP tool: the preview is the message exactly as the team
   receives it, and the confirm sends it once; `escalation_read` lists
   your organization's own requests (status `requested`; there is no
   acknowledged state, the team answers at the contact you gave);
   `context_read` returns the escalation rules and the notes the platform
   team wrote for your organization. Read `context_read` before you
   escalate.

   The skill collection is `skill_read`: instruction sets Marshall follows
   for one task, the global collection plus your organization's own
   skills. Each index entry carries `default`. A platform skill marked
   default is installed for every member of every organization, holds
   no install slot and reads `installed: true` in the index;
   `nc-request-invoice` is the first such skill. Read a skill's body
   before following it. A skill never overrides the market rules or an
   approval gate above. `skill_write` creates, updates and deletes your
   organization's own skills in one call; `propose` is two phase and
   asks the team to add one to the global collection.

   Telling the platform team how the platform worked for your human —
   the cluster, scheduling, pricing, the docs, these tools, billing — is
   the `feedback_write` MCP tool: one call, recorded for the team, no
   reply (a question that needs a human answer is `escalate`).
   Public-research organizations cannot earn feedback rewards, even after
   buying credits; their feedback is still recorded. For eligible standard
   organizations, substantive feedback, 500 or more characters after whitespace
   collapse, earns the organization 25 credits at submit time, at most
   once per member per rolling week and at most once per $1,000 of
   credits your organization has loaded or received (lifetime; rewards
   themselves do not count). Shorter feedback is recorded and read, and
   earns nothing. The response says whether the reward was granted and,
   if not, why; `feedback_read` lists your organization's submissions
   and your own eligibility before you write. Rewards exist to hear
   real experience: padded, duplicated or generated submissions make an
   organization ineligible for rewards at National Compute's sole
   discretion. Write what your human actually experienced, once.

   Publishing a finished thing — an html report, a markdown write-up, an
   image, a dataset sample, model weights, any file — for your human's
   organization or for everyone is a **checkpoint**. Inside a workspace
   Marshall publishes with `checkpoint_publish` (a workspace file; the
   organization by default — the public feed is National Compute's own
   organization's for now and any other organization's public publish is
   refused before a confirmation; your human confirms every publish, a
   public one twice) and opens one
   in a pane with `checkpoint_open`. `checkpoint_read` lists and searches
   your organization's checkpoints; `checkpoint_write
   delete` removes one of your own in two phases. A public item appears
   on the feed only after National Compute has reviewed it, under the
   organization's name and the publisher's username. Never publish
   anything that looks like a credential or token: every door refuses it,
   and a refusal is final. Publish only when your human asks you to
   share or publish something, never on your own initiative.

## Memory

Marshall is the platform's own agent. Only a Marshall workspace (the
console's Marshall page, a daemon on a machine dedicated to the member
with shell panes and a disk that survives restarts) has memory. The
classic console chat, the Marshall terminal program and MCP clients
hold no notes. For them organization context comes from the platform:
`context_read source=org_context` returns the notes the platform team
wrote for your organization. Skills come through `skill_read`.

A workspace agent has two trees of notes. The private tree is
`/data/memory/<slug>.md` on the member's workspace disk; only that
member's Marshall reads it. The organization tree is
`/org/memory/<username>/<slug>.md` in the organization folder every
member mounts; every member's Marshall reads the whole tree. Write
only into your own folder there. Nothing on disk enforces the folder
name. Marshall rebuilds the note index from disk before every turn and
lists the newest notes up to a cap; `memory_search` finds the rest and
`memory_read` opens one. The tools address a private note by its slug
and an organization note as `org/<username>/<slug>`. `memory_save`
writes the private tree without asking. `memory_save root:"org"`
promotes a note to the organization tree; while Full Access is off (the
default is stated above) it opens a card naming the exact path. Share
a note only when your human asks for it to be shared. The workspace
writes the live paths, the gate, the work log path under
`/data/worklog/` and the no secret rule to `$HOME/AGENTS.md` at every
start for an agent started in a shell pane.

A shell agent may write a note file into the private tree by hand; the
next turn indexes it. A note is one markdown file with a frontmatter
block between two `---` lines. The slug is lowercase letters, digits
and dashes, at most 80 characters. Required keys: `name` (the slug),
`description` (one line of at most 200 characters), `author`,
`created` and `updated` (ISO 8601 dates). Optional: `type` (`fact`,
`decision`, `preference`, `procedure` or `reference`; `fact` when
absent); `supersedes: <ref>` hides the older note from the index when
the superseding note is the newer of the two and both sit in the same
tree (the hidden note stays readable); `expires: <ISO date>` drops the
note from the index from 00:00 UTC of the stated date. The index skips
a note that lacks a required key or exceeds 4 MiB. A note never holds a secret or a token. `memory_save`
refuses a save that looks like a credential and writes nothing; a file
written by hand gets no such check. Every note Marshall reads back
passes through redaction. The secret still sits on disk.

A single member organization has no organization folder and no
organization tree. Both appear when the organization gains a second
member and moves to a hub of its own. The member's workspace disk moves
whole and private notes move with it.

## The market mechanism

1. **Declare, don't reserve.** For a VM cluster, `PUT /api/vm/capacity`
   with `max_gpus` (how many GPUs you want — always a multiple of the
   payload's `gpus_per_node`, because capacity is granted as whole
   nodes; off-grid values are refused, never rounded) and
   `max_price_per_gpu_hour` (your ceiling, USD, whole cents). For a
   Kubernetes cluster there is even less to say: the GPU jobs you
   apply create the demand, and `PUT /api/k8s/bid` sets your limit
   price, the most you will pay per GPU hour. With none set the
   cluster bids $0, so the demand exists but loses every tick until
   you set one.
2. **The market clears in whole nodes.** Short ticks re-run a
   sealed-bid combinatorial auction over all demand. You are billed
   what the auction assesses for capacity you hold — second-price in
   nature, never above your ceiling, usually below it. Rates are per
   cluster, not uniform, so there is no single posted clearing price.
3. **Prices come back to you.** `GET /api/market/ticks` publishes
   clearing statistics per tick — see "Reading the order book" for the
   units and the traps. Capacity and limit price reads echo your declared
   ceiling. A declare too low to ever clear the market is refused
   outright (`bid-too-low`); a ceiling above $100/GPU-hour too
   (`price-above-maximum`). Market conditions can also move over a
   standing declare (the write-time check never re-runs): the VM
   capacity GET then reports `declared.bid_too_low: true` — that ask
   can never fill until you raise the ceiling. The k8s equivalent is
   `bid_too_low` on the limit price read (`GET /api/k8s/bid`). What
   to bid is read off the market's own numbers — the ticks and the
   price to win ladder below — never off a posted price.
4. **Scaling down is always free.** Lower `max_gpus` (VM) or delete
   jobs / set the limit price to `0` (Kubernetes) any time; only what
   you acquire is priced.

## Reading the order book

- `GET /api/market/ticks?limit=50` — one row per auction tick, newest
  first. `price` is the volume-weighted mean of assessed rates in
  **USD per NODE-hour**: divide by the payload's `gpus_per_node`
  before comparing it to a per-GPU-hour ceiling. `price_min` /
  `price_max` give the spread across winners that tick.
- `GET /api/market/history?hours=168` — the bucketed trend, up to a
  year back.
- `GET /api/k8s/market/ladder` — for a Kubernetes cluster, the price to
  win ladder: per job size in whole nodes, the minimum bid that won that
  many nodes at each clearing, and `now` for the newest tick — the
  honest entry cost for a job of k nodes, in the same USD per NODE-hour
  unit as the ticks. `GET /api/k8s/market/spend` is what your own nodes
  are assessed this hour; `/paid` is that over time; `/demand` is your
  cluster's GPUs by band (bidding, provisioning, scheduled, reserved).
  All resolve the cluster like the limit price read (binding or
  `?cluster=`).
- `price: null` means nothing cleared that tick — no demand beat the
  reserve, or there was no supply. `0.0` is a real rate.
- `enabled: false` on the payload means the feed is not armed for this
  market; the empty data tells you nothing either way.
- A tick where `price_min == price_max` usually reads as one bidder
  (or perfectly uniform contention); a widening spread means competing
  ceilings are being assessed differently.
- What you are assessed is capped by your own ceiling — so a rate
  pinned exactly at your ceiling usually means the true displaced
  value sits at or above it, not that the market happens to equal it.

## Setting a limit price that actually clears

Nodes sell whole: capacity is priced and granted a node at a time.

On the Kubernetes surface each request bids its gang's GPU total ×
your limit price — the whole job, priced together. Admission enforces
the shape that clears: every pod template must request a multiple of
`gpus_per_node` GPUs — full nodes only, a multi-node job as several
full-node pods (a partial-node submit is refused with
`JobSizeTooSmall`; templates requesting no GPUs are exempt). The
exception is a cluster with reservation packing enabled. There a pod
may request fewer GPUs than one node. Such a pod runs only on your
reservation's nodes. It never bids on the market. A pod larger than one
node must still request a multiple of `gpus_per_node` GPUs. The packing
cluster refuses other shapes with `JobSizeNotNodeAligned`. When packing
is reverted on a cluster, new sub node pods, Deployment updates and Job
retries are refused with `JobSizeTooSmall` until packing is switched on
again for the cluster. Resubmit that work as whole node pods. When a
reservation ends, a running sub node job is requeued. It waits until
reserved nodes exist again.

What to bid is read off the market's own numbers, never a posted
price: the price to win ladder (`GET /api/k8s/market/ladder`) quotes
the entry cost per job size right now, and the ticks carry the recent
clearing range. A ceiling too low to ever clear the market is refused
at write time (`422 bid-too-low`) — pending demand is free, but a bid
that can never win means nothing will ever arrive. Market conditions
can also move past a standing bid: the limit price read then echoes
`bid_too_low: true`, and the market posts the verdict inside the
cluster, on the job itself — a priced-out request reads `BidTooLow`
as an event on the Job, clearing the moment the bid becomes viable —
the operating loop below has the full reason list. Watch those
instead of polling ticks.

One waiting verdict is NOT about your price: when the site itself is
out of capacity of the class your workload needs, no bid clears it.
Raising yours will not start the job sooner, so do not spend on it.
Read every waiting verdict from `GET /api/k8s/workloads`; the console
shows nothing this read does not. `services[].capacity_unavailable`
and `services[].bid_too_low` are the two verdicts as booleans per
Service. `pending[].reason` carries the market's verdict when the
market has one, always one of these eight strings: `outbid` (with
`win_price_per_gpu_hour` when the market quoted one), `waiting for
available supply`, `cluster too small for this request`, `waiting for
capacity; none available at this site right now`, `bid too low to
ever clear`, `node provisioning`, `waiting for reserved GPUs`,
`reserved GPUs too few for this request`. The two reserved verdicts
appear only on clusters with reservation packing enabled. They name a
request smaller than one node. Never raise the limit price on either.
The request runs only on your reservation. `waiting for reserved GPUs`
clears when reserved GPUs free. `reserved GPUs too few for this
request` names a job whose pods are smaller than one node while the
job spans more than one node: submit whole node pods to run across
nodes. Or submit fewer pods so the job fits one node. Any other value
is scheduler state: the scheduler's own reason for a pod the market
is not judging (`Unschedulable`, `ContainerCreating`,
`ImagePullBackOff`, a static cluster's words), the request's condition
message, or empty before the first verdict. Branch on the two booleans
and the eight verdicts, and treat everything else as scheduler state to
watch with kubectl.
`services[].requests[].reason` is the market's full sentence with the
numbers. `market.price_to_win` maps a job size in nodes to the price
per GPU hour that wins it right now.
In cluster the request's reason reads `PendingSupply`, `Outbid` or
`Pending` as usual, and the market feed's ticks read `null` while
nothing clears. The platform retries on its own and the job starts
when capacity returns. The VM surface reports the shortage directly:
`declared.capacity_unavailable: true` on the capacity GET.

Strategy, given second-price assessment:

- **Set your true limit price, once.** A limit price above the going
  rate cannot raise what you pay — you pay for displaced demand, never
  above your ceiling.
- **Do not probe upward in small steps.** Raising the ceiling also
  raises the assessment cap on nodes you ALREADY hold; contention your
  old ceiling was hiding re-prices them at the new cap the very next
  tick. Several small raises cost strictly more than one honest one.
- Winning contested nodes takes them from someone, and the same can be
  done to you — but never inside the minimum-duration protection
  window a fresh grant carries (currently two hours for a Kubernetes
  node, one hour for a VM node): while it is open, no competing bid
  can take the node. On Kubernetes the window rides your live
  ceiling — cutting it below the price the job held when its gang
  started voids the protection at once; raising it never does. When
  the window lapses, the node competes normally again, every tick.
- **Don't slice work into super-short jobs.** The protection window
  is also the minimum sensible unit of work: while it is open a node
  can be assessed at up to your full ceiling, and on Kubernetes the
  first hour of a grant is a minimum hold — the node stays yours, and
  bills, for at least an hour even if the job that asked for it
  finishes sooner. A ten-minute hold pays peak rate for a full hour
  on top of provisioning overhead it never amortizes, and every short
  burst re-pays node join, image pull, and the idle-shed grace after
  it exits. Batch small tasks into holds of at least an hour rather
  than grabbing and releasing nodes per task.

## The operating loop (Kubernetes)

1. Deploy. The GPU jobs you apply **are** the ask — there is no
   node-count API. Apply a `Job`, `JobSet` or `MPIJob` whose pods
   request full-node GPU multiples: Kueue, installed in the cluster,
   creates it suspended and files one request for the whole gang —
   every node the job needs, priced together, granted together.
   Nothing else is needed — no queue name, no `suspend: true`, and
   the GPU-taint toleration is injected on admission. (A `Deployment`
   replica or a bare GPU pod is its own one-node request, held
   `SchedulingGated` until granted; CPU-only pods run at once on the
   cluster's CPU worker and neither bid nor hold GPU nodes.) Submits
   are admission-gated: full-node GPU shapes only (`JobSizeTooSmall`)
   and a balance covering two hours at your ceiling (`BalanceTooLow`)
   — both are detailed above. On a cluster with reservation packing
   enabled a pod smaller than one node is admitted too. It runs on your
   reservation.
2. Set the limit price, then read the queue on the job — never on pods. A
   waiting job has **no pods**: `kubectl --context <cluster> get pods` stays empty until
   the whole gang is granted, so an empty pod list means a queued
   job, not a broken cluster. `kubectl --context <cluster> get workloads` lists what
   waits; `kubectl --context <cluster> describe job <name>` carries the market's verdicts
   as events, one per change of reason — `BalanceTooLow`,
   `BidTooLow`, `NoSupply`, `PendingSupply`, `Outbid`, `Pending`,
   `ReservedBusy`, `ReservedTooSmall` (a fresh request reads nothing
   until the market's next tick; the two reserved events appear only
   where reservation packing is enabled) —
   then `NodeGranted` once per granted node. Nodes provision and join
   in minutes, the gang's pods all start in the same moment, and
   first start still pays image pull and any model download. A job
   can also wait before any request exists: beyond the cluster's max
   size (the quota of Kueue ClusterQueue `market` — platform-owned,
   edits refused with `StationOwned`) it waits in your own queue
   under `QuotaReserved`, where no ceiling helps. The
   `kueue.x-k8s.io/priority-class` label (`low`/`normal`/`high`, set
   on the suspended Job) orders your own jobs while that cap binds;
   it never buys market position over another tenant.
3. Reclaims are part of the loop, not failures. When the market
   clears above your ceiling, the whole gang is suspended, its pods
   deleted together, and the job requeues as a new request at your
   current ceiling — without consuming the Job's `backoffLimit`. An
   outbid reclaim gives **one minute's warning**: the request gains
   `PreemptionNotice=True`, `NodePreempting` events land on the node
   and its pods, every pod gets a
   `marketplace.nationalcompute.com/preempt-at` annotation naming
   the deadline (readable in-pod through a downward-API volume), and
   eviction delivers SIGTERM with the window as its grace — SIGTERM
   is your checkpoint signal, `preempt-at` the deadline to checkpoint
   by. A reclaim your own account causes (a limit price cut or withdrawn, the
   balance running out) skips the warning: the drain starts at once,
   with the same 60-second eviction grace. Set
   `terminationGracePeriodSeconds: 60` so pod grace matches the
   window, and make jobs resumable.
4. Checkpoint to the shared volume, not the node. Node-local disk is
   ephemeral — a reclaimed or shed node is destroyed, caches
   included. Clusters provisioned with shared storage mount one
   shared volume at the same path on every node (`/mnt/shared`
   today; reach it with a `hostPath` volume): the place for datasets,
   checkpoints and outputs that must outlive a node.
5. Idle nodes shed after the idle grace — once past the first hour's
   minimum hold — and billing stops. An unfilled standing ask costs
   nothing — leave it in place and capacity arrives whenever the
   market allows.

Two reads back the loop with the console's own verdicts. Storage
fullness: `GET /api/k8s/storage` (MCP: `grid_api get`) returns the
Storage page for every cluster of your organization, and every
fullness reading carries `level`: `ok`, `warn` (above 90 %), `crit`
(at or above 99 %), or `null` when unknown; `thresholds` names the
constants. Read `shared[].level` before a checkpoint fills the shared
volume, `nvme.rows[].level` for node scratch, `quota_level` for the
shared home. GPU health over time: `GET /api/k8s/capacity/history`
(MCP: the same path through `grid_api get`) returns GPUs by health
state per bucket, `states[].label` one of seven words: `healthy`,
`degraded` (still schedulable), `unschedulable`, `cordoned` (your own
admin's park, still billed), `moving`, `unreachable`, `unknown`; plus
`scheduled`, `demand` and `available` for demand against capacity.
Nothing on these reads is a market verdict; a shrinking `available`
with `demand` above it is the signal to look at the limit price.

## Ground rules

- Auth: `Authorization: Bearer <token>`. A cluster-bound token needs
  no cluster name; an org-wide token over several clusters of a kind
  must pass `cluster` (query param on GETs, body field on writes) or
  the call fails `409 ambiguous-cluster` — whose payload lists the
  candidate `clusters` by name. `GET /api/whoami` gives you the same
  list up front, with kinds.
- kubectl: ALWAYS name the cluster explicitly on every invocation —
  `kubectl --context <cluster> …`, the context the setup script
  merges. Never rely on the ambient current-context: your human's
  default may point at a different cluster entirely, and a
  right-command-wrong-cluster mistake looks like success. Sign in
  recurs: after about a day without kubectl use, the next call prints
  a fresh sign in link and waits for the browser step. Run kubectl
  where you can read its output while it runs; a call that seems hung
  is showing a link — relay it to your human, as at setup.
- Read `gpus_per_node` from the payload; never hardcode a conversion.
- Capacity writes, the VM declare and the Kubernetes limit price alike: echo
  the payload's `version` back as `expected_version`. The write applies
  only while it still matches. A `409 version-mismatch` echoes the live
  version; re-read and decide again. The version counts every accepted
  write, a withdrawal included, and never restarts. Read `updated_by`
  on the limit price read first: `token:<id>` is an agent, anything else is a
  human's decision.
  `GET /api/org/activity` shows every write on the organization, human
  or token, newest first.
- Writes are rate limited per cluster (the spec carries the cadence).
  Every token also has a budget across all routes, 300 requests per
  minute by default; the vm contract states the figure in force. The
  call past either limit answers `429 rate-limited` with
  `retry_after_s` and a `Retry-After` header. Space your polls: the
  price feed and the limit price read change once per tick, so one read every
  10 s is plenty. Errors are plain JSON `{"error": "<code>", ...}` per
  the spec's error catalog.
- VM nodes are identified by public IP everywhere on this API.
