Coyos service

The qualification process for one AI task.

Coyos does the work below with you, for one task and one deployment. You receive the files. After that you run the tests yourself. This is not a licence for the proposed packs, and it is not the demonstration pack.

What this is, and what it is not

Three different things. The public fixture coyos.fs.contact-routing@0.3.1 is an example you can run. The commercial pack is coyos.fs.contact-routing at the version Coyos delivers, and that is the pack you can buy for a routing task. This service is neither: it is paid work to establish the process for one task that is not that pack. Proposed packs stay unavailable. A QUALIFIED result does not grant permission to use the model without human review.

Work Coyos does

Establish the process. Run it once with you.

The engagement is the task you name. It is not a copy of the public fixture, and it is not a second name for the commercial contact-routing pack.

  1. 1

    Name the task and the deployment the tests apply to: what the model is supposed to do, and the endpoint, prompt, and serving settings the run must record.

  2. 2

    Define expected results with you, from cases you supply.

  3. 3

    Identify which errors are critical, and score those from reference checks rather than from a judge alone.

  4. 4

    Separate the corpora the runner already names: qualification cases the decision reads, held-out calibration cases, and challenge cases reported beside that decision.

  5. 5

    Select acceptance limits and the number of test cases those limits need. The number is written down with the limit.

  6. 6

    Assemble a pack directory the runner can validate.

  7. 7

    Run the tests once with you against the deployment you name, when that endpoint is reachable from your network.

  8. 8

    Walk through failures and through a result that does not support a decision: NOT_QUALIFIED and INDETERMINATE.

  9. 9

    Show the commands to repeat the tests after a deployment change. Those commands are the runner’s, not a Coyos-only tool.

What you supply

Your task, your deployment, your cases.

  • The task: what the model must do, and what it must not do.
  • The deployment: endpoint URL, model identity, system prompt, and serving settings the fingerprint must record.
  • Your own test cases, with expected results. Etalon does not invent your acceptance set.
  • Which errors you treat as critical.
  • The acceptance criteria you want the pack to enforce: limits, and the case counts those limits require.
  • A person who can run the Etalon runner on your network against that endpoint.

Files you receive

A pack you can validate, and a bundle if a run finished.

File names inside the pack follow the runner. See Packs in the docs. This page does not invent a second layout.

A versioned pack directory the runner can validate. The packs page in the docs names the files: pack.yaml, corpus/qualification.jsonl, corpus/calibration.jsonl, corpus/challenge.jsonl, rubrics/, evaluators.yaml, thresholds.yaml, qualification.yaml, triggers.yaml, and methodology.md.

Limits and case counts written into thresholds.yaml and qualification.yaml, with the reason for that count in methodology.md.

The evidence bundle from the worked run, when a run is completed during the engagement.

The qualify, inspect, verify, and report commands for your pack path, the same verbs as the getting-started page.

Without Coyos afterward

You repeat the tests.

The runner is the thing that runs the pack. Coyos is not in that path after delivery.

You can do these without us

Run the pack again with etalon qualify against the same deployment.

Verify the bundle with etalon verify on a machine that has the runner and the bundle.

Repeat the run after a deployment change, and compare the new bundle with the previous one.

Read NOT_QUALIFIED and INDETERMINATE from the bundle. Those results do not support a decision to treat the task as cleared.

Install and command names stay the ones on getting started. You point them at your pack directory and your endpoint. A QUALIFIED result does not grant permission to use the model without human review.

One task. Your cases. Files you keep.

Write with the task, the deployment, and whether you already have labelled cases. You will get a plain answer about whether this service fits.