Evidence Logged experiments6 October 2026
Every decision,
accounted for.
A real local encoder. 3,080 labelled banking messages. The accepted decisions, the ones sent for review, and the mistakes—all from the same frozen test run.
Classify before Qwen. Measure both routes.
SDK 0.3.1 · local reference70.5% fewer LLM calls. Accuracy changed from 69.5% to 83.4%, below our preset 90% floor. This configuration did not pass.
The frozen sample contains 770 public BANKING77 messages, ten per category. Direct Qwen classifies every message; the hybrid runs MemCat first and makes a fresh Qwen request on review or failure. Both return an exact label or review. These are classification calls; no answers are generated for a customer.
| Measured outcome | Qwen directly | MemCat + Qwen |
|---|
| Correct / 770 | 535 / 770 (69.48%) | 642 / 770 (83.38%) |
|---|
| Wrong / review / failed | 218 / 17 / 0 | 123 / 5 / 0 |
|---|
| Actual LLM calls | 770 | 227 |
|---|
| Reported prompt + output tokens | 2,103,931 | 620,541 |
|---|
| Sum of measured query times | 299.67 s | 87.72 s |
|---|
| Per-message p50 / p95 | 353.57 ms / 794.05 ms | 8.41 ms / 436.41 ms |
|---|
MemCat accepted 543 messages locally: 527 were correct and 16 were wrong (97.05% accepted precision). On the remaining 227, the actual Qwen fallback returned 115 correct labels (50.66%), 107 wrong labels and 5 reviews. Most final errors came from the fallback route.
The paired accuracy difference was 13.90 percentage points; the one-sided 95% lower bound was 11.95 points. The interval uses 20,000 predeclared stratified paired bootstrap samples. Review and failure count as incorrect on this in-scope sample. This is an internal tolerance, not customer approval.
| Target frozen before inference | Result |
|---|
| All 770 cases recorded on both routes | Met |
|---|
| At least 90% correct across all hybrid inputs | Not met |
|---|
| Paired 95% lower bound no worse than −2 percentage points | Met |
|---|
| At most 1% hybrid operational failures | Met |
|---|
| At least 50% fewer actual LLM requests | Met |
|---|
| At least 50% fewer returned prompt + output tokens | Met |
|---|
| At least 30% less measured query time | Met |
|---|
Qwen2.5 7B Q4_K_M ran locally through Ollama 0.35.1 on an Apple M4 Max with an 8,192-token context. It received the 77 labels, descriptions and one training example per category. MemCat used 32 examples plus a description per category. This comparison does not establish the best possible Qwen configuration.
Category descriptions were literal names. The upstream label get_physical_card is attached to PIN-related examples, which can mislead a prompted model. An independent source check verified every imported label against the original CSV files. The frozen prompt and this failed result are preserved; future prompt changes require separate evaluation.
Qwen ran first, then the hybrid, with the target model unloaded and the same three excluded warmups repeated between routes. Prefix caching remained active. Query time includes local classification and fallback when used; setup, warmups and receipt writing are excluded and recorded separately. Fixed route order may affect timing. The sample was previously evaluated by MemCat; it is not a fresh held-out test or customer traffic. Known mixed-intent failures below remain unresolved.
Comparison report · SHA-256582a276a3995f5c10be8553b632fe60a9b82462a108fdaeb0bde82f5c5dbd858
Public BANKING77 data: Casanueva et al. / PolyAI, CC BY 4.0. Dollar savings, total compute cost and customer adoption remain unmeasured. Sources and licenses.
What happened to the messages?
All 3,080 official BANKING77 test rows, across 77 categories. Thresholds and prototypes were frozen before this test. No test rows were removed.
English · banking supportCorrect among automatic decisions97.48%2,092 correct / 2,146 accepted
Automatically classified69.68%2,146 accepted / 3,080 total
Incorrect automatic routes541.75% of all 3,080 inputs
2,092 correct54 incorrect934 review0 failed or skipped
95% Wilson interval for correctness among accepted decisions: 96.73%–98.07%. Review means the classifier declined to choose; no fallback model ran, and those messages are not counted as correct answers.
Precision and coverage belong together.Without abstention, its top-choice accuracy is 88.25% and macro-F1 is 0.8821. The higher accepted precision comes from sending harder messages to review.
Compared with simpler options.
The same test inputs and labels. Each baseline used only training data; rejection thresholds were selected on calibration data. The local text baselines require no neural encoder.
| Method | Top-choice accuracy | Accepted precision | Coverage | Wrong routes |
|---|
| MemCat + MiniLMExample similarity + abstention | 88.25% | 97.48% | 69.68% | 54 |
| Exact training lookup | 0.13% | — | 0.00% | 0 |
| Category word overlap | 36.66% | — | 0.00% | 0 |
| TF-IDF nearest centroid | 81.17% | 96.93% | 46.46% | 44 |
| Always review | — | — | 0.00% | 0 |
“—” means no automatic decisions, so precision is undefined. Exact lookup and category overlap did not meet the calibration target and accepted no test messages under their selected rejection policy. Always choosing the most common training label achieved 1.30% accuracy.
Measured speed. Defined boundaries.
Individual, sequential queries on Apple M4 Max, Node 22.22.3. These measurements include token validation, tokenization, q8 encoding, category matching, and the review decision.
- Warm latency · median
- 1.78 msPer message; 3 warmup calls, no result cache
- Warm latency · p95 / p99
- 2.70 ms / 3.72 ms
- Model load from local disk
- 371.73 msExcludes first network download; not a browser cold-start figure
- Sequential throughput
- 524 messages / secondMeasured across the full test callback loop
- Example index
- 10.07 MB JSON / 3.67 MB gzip2,537 prototype vectors; encoder runtime and weights are additional
- Decision policy
- Similarity ≥ 0.70 · margin ≥ 0.08Cosine scores are not probabilities of correctness
Runtime: ONNX CPU; individual sequential queries. Browser WASM uses a different execution environment; the live demo reports its own loading and inference times. These figures do not measure an LLM workflow’s end-to-end latency or cost savings.
Run the same policy in your own project.
The Node starter includes the encoder adapter, pinned dependencies, and a replay command. We verified a fresh installation outside this repository using the published SDK: it downloaded the assets, ran local inference, and wrote inspectable results without an API key or source edits.
This is an internal integration check on Node 22 and Apple silicon. Its four authored examples are separate from the accuracy benchmark above. Setup time, first inference, asset requests, and failures are logged separately; no LLM cost savings were measured.
Where it still struggles.
Passing an in-scope benchmark does not make every decision safe. We also ran a separate unknown-topic diagnostic and an authored stress suite.
Narrow diagnosticDistant, non-banking requests
0 / 240Accepted as a banking category. All official CLINC test rows from eight predeclared tasks, including weather, recipes, and music. The 95% Wilson upper bound is 1.58%.
This does not establish rejection of unfamiliar requests close to banking topics.
Stress gate failedMixed intent, negation, and spelling
3 / 14Authored cases missed their expected handling. One mixed request was routed automatically; two other cases were deferred instead of receiving the expected category.
These are regression cases, not a representative accuracy sample.
- Wrong automatic decision
“Why was I charged twice, and how do I close my account?”
- Expected
Review- Returned
transaction_charged_twice- Case
- mixed-02
- Deferred · expected a category
“I do not want to close my account. Where is my card?”
- Expected
card_arrival- Returned
Review- Case
- negation-02
- Deferred · expected a category
“my caard got stoln”
- Expected
lost_or_stolen_card- Returned
Review- Case
- typo-01
The incorrect automatic decisions.
The first 10 of 54 errors in source order. Related banking categories are sometimes confused. The full report contains every result, expected label, similarity decision, and timing.
“Is there a way to know when my card will arrive?”
- Expected
card_arrival- Returned
card_delivery_estimate- Source ID
- banking77-test-00004
“How long does a card delivery take?”
- Expected
card_arrival- Returned
card_delivery_estimate- Source ID
- banking77-test-00012
“How do I change currencies to euros?”
- Expected
fiat_currency_support- Returned
exchange_via_app- Source ID
- banking77-test-00256
“Can I get my card expedited?”
- Expected
card_delivery_estimate- Returned
card_arrival- Source ID
- banking77-test-00309
“How do I unblock my card using the app?”
- Expected
card_not_working- Returned
pin_blocked- Source ID
- banking77-test-00369
“Can I change from one currency to another?”
- Expected
exchange_via_app- Returned
fiat_currency_support- Source ID
- banking77-test-00403
“What all currencies can be exchanged?”
- Expected
exchange_via_app- Returned
fiat_currency_support- Source ID
- banking77-test-00411
“Is it possible to change to another currency?”
- Expected
exchange_via_app- Returned
fiat_currency_support- Source ID
- banking77-test-00415
“Can I change to another currency?”
- Expected
exchange_via_app- Returned
fiat_currency_support- Source ID
- banking77-test-00419
“Change currency”
- Expected
exchange_via_app- Returned
fiat_currency_support- Source ID
- banking77-test-00426
What this result does—and doesn’t—prove.
One declared workflow
English, single-intent banking messages with 77 published labels. Public benchmark performance does not establish accuracy on your support queue, language, or category definitions.
Existing source overlap, disclosed
7 official test messages match earlier source rows after normalization. All test rows are retained. Excluding those overlaps gives 97.48% accepted precision at 69.61% coverage across 3,073 messages. Public data may also have appeared in encoder pretraining.
An existing encoder, with controls
MemCat uses MiniLM embeddings and nearest-example cosine similarity. It adds explicit thresholds, serializable indexes, evaluation, tracing, and shadow mode. This is not a newly trained foundation model.
Economics still need your workload
This local run made 0 hosted classification calls. Local compute and review work still have costs. The separate local Qwen comparison above measures real classification requests and latency. Customer savings, paid-provider bills and independent adoption remain unmeasured.
Keep the receipts.
Training, development, and calibration contained 7,963, 965, and 1,071 rows respectively. The selected model and thresholds were frozen before testing. The report records dataset hashes, source-code hashes, outcomes, baselines, and limitations.
Full report · SHA-2568d29ebcbfa71e7000b095c983fe74b5cc4e133cd4c8aaacf19844ead5e115c41
BANKING77: Casanueva et al. / PolyAI, CC BY 4.0. CLINC150: Larson et al. / Clinc, CC BY 3.0. MiniLM encoder: Apache 2.0. Full attribution and immutable source revisions are in the linked notice.