# MemCat 0.3.1 — quickstart

A private downloadable alpha for example-based classification, evaluation, shadow comparison and local/model routing. Not published on npm; UNLICENSED.
Requires Node.js 22+. The SDK has zero runtime dependencies; your embedding runtime and model weights are separate.

Download https://usememcat.vercel.app/downloads/memcat-sdk-0.3.1.tgz into your project, then:

```sh
npm install ./memcat-sdk-0.3.1.tgz
```

## Start with a measured replay

### Run the real semantic demo in Node

Download and extract the [Node starter](https://usememcat.vercel.app/downloads/memcat-semantic-starter-0.1.1.tgz) into a new folder. Its [archive checksum](https://usememcat.vercel.app/downloads/memcat-semantic-starter-0.1.1.sha256) identifies the exact release. In the extracted `package` folder:

```sh
npm ci
node run.mjs --input "The cash machine kept my card. What do I do?" --output ./message-receipt.json
node run.mjs --dataset ./examples.json --output ./examples-run.json
```

This runs the same frozen 77-category English banking policy as the browser demo using the published SDK, a pinned MiniLM model and an included encoder adapter. No API key or code edits are required. First use downloads a roughly 10 MB index, 23 MB of model weights and supporting assets; runtime dependencies are additional. Setup and inference timings are recorded separately. Read the included README for provenance, cancellation, text privacy and limitations.

The authored examples test integration, not new accuracy. The banking policy still makes mistakes; it is not a ready-made classifier for arbitrary categories. Use your own labelled cases before changing routing. Local compute is not free, and no LLM savings comparison is claimed.

### Verify just the SDK evaluator

The package includes a working evaluator and shadow-mode pilot. Verify your installation without keys:

```sh
node node_modules/@memcat/sdk/examples/replay.mjs \
  --module ./node_modules/@memcat/sdk/examples/fixture-classifier.mjs \
  --dataset ./node_modules/@memcat/sdk/examples/fixture-cases.json \
  --output ./memcat-replay-001.json
```

This fixture uses two literal matches; it verifies integration, not semantic accuracy. The immutable receipt records coverage, mistakes, deferrals and timing. Unknown costs remain null. Use a new output filename on each run.

Read `node_modules/@memcat/sdk/PILOT.md` to replace the fixture with your current classifier and a MemCat candidate, supply human-labelled examples, then run shadow comparison without changing your application's result. Raw input text is omitted by default.

For semantic classification, `createClassifier` from `@memcat/sdk/classifier` accepts category descriptions/examples, an embedding adapter, calibrated similarity/margin thresholds, and bounded callback limits. It returns a category, review, or an explicit failure. The packaged `README.md` documents the complete encoder and index contract.

Try the [real local demo](https://usememcat.vercel.app/classify) first. Its bank-specific index uses MiniLM; your categories require their own examples and evaluation. Inspect the [measured results and known failures](https://usememcat.vercel.app/evidence). A similarity score is not confidence, and a review flag does not prove a fallback answered correctly.

## Optional: deterministic routing

The example below verifies the retained callback router and cache. It does not run a semantic model.

## 1. Run a local task (no key, no network request)

Save as `local.mjs`:

```js
import { createRouter } from "@memcat/sdk/router";

const router = createRouter({
  policy: {
    version: "v1",
    routes: [
      {
        id: "uppercase",
        match: { field: "context.task", op: "equals", value: "uppercase" },
        handler: "uppercase",
      },
    ],
    fallback: "model",
  },
  handlers: {
    uppercase: {
      kind: "local",
      run: ({ input }) => ({
        text: input.toUpperCase(),
      }),
    },
    model: {
      kind: "model",
      run: () => {
        throw new Error("No provider connected yet.");
      },
    },
  },
  limits: { maxCalls: 10, maxConcurrent: 1, timeoutMs: 10_000 },
  cache: { maxEntries: 16, ttlMs: 60_000 },
});

const request = {
  input: "Hello, world.",
  context: { task: "uppercase" },
  tenantId: "example-project",
};
console.log(await router.run(request)); // completed / local
console.log(await router.run(request)); // completed / cache; callback not repeated
console.log(router.attempts); // { total: 1, local: 1, model: 0 }
```

```sh
node local.mjs
```

## 2. Connect your server to OpenAI

This step makes a real, billable provider request. Set OPENAI_API_KEY and
OPENAI_MODEL in your server environment. Choose a Responses-compatible model
your account can use. Do not put the key in a browser or NEXT_PUBLIC variable.

Save as `model.mjs`:

```js
import { createRouter } from "@memcat/sdk/router";
import { createOpenAIHandler } from "@memcat/sdk/openai";

const handler = createOpenAIHandler({
  apiKey: process.env.OPENAI_API_KEY,
  model: process.env.OPENAI_MODEL,
  instructions: "Explain the requested concept in one concise paragraph.",
  maxCalls: 1,
  maxOutputTokens: 256,
});
const router = createRouter({
  policy: { version: "v1", routes: [], fallback: "model" },
  handlers: { model: handler },
  limits: { maxCalls: 1, maxConcurrent: 1, timeoutMs: 20_000 },
});
const result = await router.run({ input: "What is exact caching?" });
console.log(result);
console.log({ httpAttempts: handler.callsUsed });
```

```sh
node model.mjs
```

To combine these examples, register this handler under `model` in the first
router and retain your local routes. Enable cache only for outputs that are
safe to reuse. Model capabilities vary; an output cap includes reasoning tokens,
so a reasoning model can reach its cap before producing text. Incomplete and
refused outputs fail; they are not presented as valid answers.

## Literal router boundaries

- Ordered literal matching, not semantic understanding. The first match wins.
- Cache is off by default. This example opts in for deterministic text transformation.
- Cache scope is not authentication. Include all relevant dependency versions
  in context; cache only fully described, reusable work. Hidden callback state
  and concurrent duplicate misses are not deduplicated.
- Limits are per instance. Router maxCalls counts callbacks. The OpenAI handler
  separately caps HTTP attempts. Neither is a deployment-wide dollar budget.
- Respect the abort signal in custom handlers. A timeout cannot preempt CPU work
  or force a noncooperative callback to stop. Its concurrency slot remains held.
- Results: completed with output, skipped with a limit reason, or failed with a
  bounded reason. Invalid configuration/requests throw before any callback runs.
- Context is sent to OpenAI as user data; tenantId is not sent automatically.
- No retries, tool execution, distributed cache, or semantic Jev selector.
- The packaged README has the full API contract. This release was verified with
  deterministic provider fixtures, not live model-quality or cost measurements.


The 0.3 SDK also includes `@memcat/sdk/classifier` and `@memcat/sdk/evaluation`. See the packaged README and PILOT.md for example-based classification, replay and shadow comparison. Model weights/runtime are separate; unknown costs remain unknown. Try the [local semantic demo](https://usememcat.vercel.app/classify) and inspect the [public evidence](https://usememcat.vercel.app/evidence).
