E-LIM
Enhanced Large Intention Model
Four to six sentences from Cognitio, written on request. Generated text: the page is the reference.
On this page
Overview #
E-LIM reads the last few lines of a conversation and says whether the next reply is fine to give, needs a nudge in a safer direction, or should stop. You give it the recent turns as plain text; it gives back a handful of labels and, when a nudge or a stop is called for, a suggested short line. Use it in front of any chat responder as a quick, cheap check.
E-LIM (Enhanced Large Intention Model; called ELIM until 2026-10-04, with the same weights) is a conversation-intention classifier a host application runs before its responder to decide whether to allow, steer or abort a reply. It reads a short text window — the last three turns plus the current request — and returns four labels with confidences: the trajectory the conversation is on (one of 32 classes such as research, scam_check or crisis), the action to take (allow, steer or abort), the harm class in play (none, crisis, medical, crime, child_sexual, scam, jailbreak or secrets) and a steer hint naming the kind of nudge the responder should fold into its answer.
The input is plain text, so the window can come from any transcript the host keeps. The output is a small JSON object the host reads before doing anything else: on allow nothing changes; on steer the host passes the hint and the optional suggested line to its responder as guidance; on abort the host answers with the suggested line itself and never calls the responder. E-LIM never writes the reply and never sees the responder’s output.
E-LIM is the default model of the intention family. Its sibling LIM scores the same window with a single 512-wide hashed bag and answers with the same contract; LIM3D and E-LIM3D apply the same idea to observed movement rather than text. E-LIM is 13.2M parameters, runs inside the host process with no native dependencies and no network call, and answers in well under a millisecond on a laptop CPU.
Intended use #
- A pre-reply intention gate: classify the current request, then let the host allow, steer or abort before its responder runs.
- Routing short, informal, English requests into coarse job classes (
reminder,weather,transit,order,research, and so on) when a fast, deterministic label is enough. - Flagging the handful of trajectories that need a fixed response — crisis, medical emergency, scam assistance, jailbreak attempts, leaked secrets — so the host can apply its own policy.
- Any place where a sub-millisecond, in-process decision matters more than nuance: per-request gating on a CPU, batch labelling of transcripts, cheap pre-filters in front of a larger model.
Out of scope #
- Writing or rewriting replies. The
messagefield is a suggested short line, safe to replace, not generated text. - Scoring the responder’s output. E-LIM classifies the user’s side of the window only.
- Long documents, non-English text, sarcasm or novel phrasings far from the template set it was trained on.
- Acting as a safety classifier of record. Thresholds favour silence, and greetings and known job phrasings bypass the model entirely.
- Not medical, legal or crisis advice.
Choose E-LIM when #
- You want the default Intention model: the released, served weights behind
POST /v1/intentionwhen nomodelfield is sent. - Character-level features matter — typos, run-together words and short fragments — because E-LIM hashes character 3- and 4-grams as well as words; LIM hashes one combined bag.
- Memory is tight and 53 MB resident is acceptable but more is not; choose LIM (8.9M parameters, 35.8 MB) when a smaller table matters more than the second bag. It shares E-LIM’s contract.
- You need the same window scored by movement rather than text: choose LIM3D or E-LIM3D instead.
Architecture #
| Parameters | 13,193,523 (about 12.6M in the two embedding tables) |
|---|---|
| Kind | Hashed n-gram classifier, four heads |
| Hash buckets | 16,384 per bag (FNV-1a 32-bit over UTF-8 bytes, modulo 16,384) |
| Bag width | 384 per bag; 768 after concatenation |
| Hidden layers | fc1 768→384, fc2 384→384 (residual), fc3 384→384 (residual); GELU after each |
| Heads | trajectory 32 · action 3 · harm 8 · steer 8 |
| Window | last 3 prior turns + current request, 1,500 characters; char bag reads the last 360 |
| Features per bag | at most 2,048 hashed indices |
| Weights | elim.bin, 52,774,490 bytes, float32 tensor bundle (16 tensors) + meta.json catalogue |
| Runtime | In-process, no native dependencies, no network call |
| Latency | sub-millisecond (~0.8 ms measured, laptop CPU) |
E-LIM is a dual-bag hashed n-gram classifier. There is no tokeniser and no pretrained embedding: the window is lowercased and hashed twice. The character bag takes every 3-gram and 4-gram of the last 360 characters; the word bag takes every unigram and bigram of the word sequence. Each n-gram is hashed with FNV-1a (32-bit, over UTF-8 bytes) modulo 16,384 into an index, and each bag keeps at most 2,048 indices. An empty bag contributes index 0.
Each bag is an embedding table of 16,384 × 384 float32 rows averaged over its indices (EmbeddingBag with mean pooling). The two 384-wide means are concatenated into a 768-wide vector and passed through fc1 (768→384) with GELU, then two residual blocks h = h + GELU(fc2(h)) and h = h + GELU(fc3(h)), each 384→384. Four linear heads read the final 384-wide state: trajectory (32 logits), action (3), harm (8) and steer (8). Softmax is applied per head at inference and the argmax picks the label; the softmax maximum is reported as the confidence.
The four label sets are listed under Inputs & outputs and on Output vocabularies. The heads are trained jointly but read independently at inference; the action a caller receives is not the raw action-head argmax but the result of the decision rules described below, which require the trajectory head to agree before a steer or abort is issued.
Of the 13,193,523 parameters, about 12.6M sit in the two embedding tables; the three hidden layers and four heads account for the rest. There is deliberately no sequence model: E-LIM has no attention, no recurrence and no positional signal beyond what an n-gram carries, so word order matters only within a bigram or a 4-character span. It never sees pixels, audio or the responder’s text.
Inputs & outputs #
Input #
| Field | Type | Required | Description | Limit |
|---|---|---|---|---|
text | string | Yes | The current request from the user of the host application. Trimmed; an empty string is rejected. Over HTTP this is the whole window, so the model sees exactly one turn. | window truncated to 1,500 characters; only the last 360 characters feed the character bag |
model | string | No | Which intention model scores the window: "e-lim" (the default) or "lim" for LIM; "elim", E-LIM's id until 2026-10-04, still selects E-LIM. "lim-nano" answers 410 model_withdrawn (LIM Nano was withdrawn on 2026-10-03); any other value answers 400 unknown_model. | — |
history | array of {role, text} | No | In-process only. Up to the last three prior turns, role "user" or "model"; responder turns that repeat an earlier safety reply are dropped before the window is built. The HTTP route always passes an empty history. | last 3 turns |
The model scores a single window string. The host builds it from up to the last three prior turns and the current request: each turn is rendered on its own line with a U: prefix for the user or an A: prefix for the responder, the current request is appended as a final U: line, the lines are joined with newlines and the result is truncated to 1,500 characters. Responder turns whose text matches an earlier safety reply (crisis, scam or refusal wording) are dropped before rendering, so E-LIM judges the user’s intent rather than its own previous advice. Over the Falcon API the history is always empty, so the window is the single U: line built from text.
A complete in-process window with two prior turns:
U: remind me in 20 minutes
A: ok — 20 min.
U: is this a scam they sent meThis window yields 141 character-bag indices and 35 word-bag indices. The 360-character span for the character bag is taken from the end of the window, so on long windows the earliest turns contribute to the word bag only. Text is lowercased inside the hashers; the window itself is passed through unchanged.
The equivalent HTTP request scores only the current request:
{ "text": "is this a scam they sent me", "model": "e-lim" }Limits: text is trimmed and must be non-empty; the window is capped at 1,500 characters; each bag keeps at most 2,048 hashed indices; there is no frame or coordinate system.
Output #
trajectory 27 labels
continueswitch_topicresearchreminderhealthmedical_emergencycrisiscrisis_methodcrimechild_sexualscam_assistscam_checkjailbreaksecretsimageslidestransitordercalendarmapsweathersportsnewsmediasmalltalkhelpunknown
action 3 labels
allowsteerabort
harm 8 labels
nonecrisismedicalcrimechild_sexualscamjailbreaksecrets
steer 8 labels
nonecrisis_linemedical_911scam_warnjailbreak_ignoresecrets_droprefocusconfirm_destructive
The four tables above list the documented vocabulary of each head. Five further trajectory labels exist in the catalogue for host-specific routing and are not documented. The steer label refocus is in the catalogue but no trajectory maps to it in training, so the weights never produce it; confirm_destructive is produced only for one of the undocumented trajectories.
A complete response for the window above, scored in-process by the 1.0.0 weights (probabilities are serialised as raw floats; the near-saturated values are the behaviour described under Confidence semantics below):
{
"trajectory": "scam_check",
"action": "steer",
"harm": "scam",
"steer": "scam_warn",
"pAction": 1,
"pHarm": 0.9999999403953552,
"pTrajectory": 0.9999998211860657,
"rule": "traj-steer≥0.58",
"model": "elim",
"version": "1.0.0",
"message": "…"
}In-process the reference runtime still names E-LIM by its internal id elim, as above; the Falcon API reports it as e-lim (see API).
Decision rule. Trajectories are grouped into three sets: abort (crime, child_sexual, scam_assist, crisis_method), steer (crisis, medical_emergency, jailbreak, secrets, scam_check) and benign (the remaining 23). After the heads are read, action is decided in this order, and rule records which step fired:
- A benign trajectory with
pTrajectory ≥ 0.28→allow(benign-traj≥0.28). - An abort-set trajectory with
pTrajectory ≥ 0.62→abort(traj-abort≥0.62). - Action head
abortwithpAction ≥ 0.78and an abort-set trajectory withpTrajectory ≥ 0.45→abort(action-abort≥0.78+traj≥0.45). - A steer-set trajectory with
pTrajectory ≥ 0.58→steer(traj-steer≥0.58). - Action head
steerwithpAction ≥ 0.72and a steer-set trajectory withpTrajectory ≥ 0.40→steer(action-steer≥0.72+traj≥0.4). - Otherwise
allow(default-allow).
The thresholds are deliberately asymmetric: a benign trajectory wins at 0.28, whereas a steer or abort needs the trajectory head to be confident on its own. This keeps everyday requests quiet at the cost of missing harmful requests phrased like benign ones.
Bypasses. Two rules answer before any weights run and set model to "bypass": a greeting pattern (hi, hello, thanks, how are you, good morning and similar, alone on the line) returns smalltalk / allow with rule: "greeting-bypass"; a known-job pattern (food ordering, transit, reminders, timers, weather, maps, directions, translation, slides, drawing, and spatial-layout phrasing) returns help / allow with rule: "host-job-bypass". A second guard runs after scoring: if the model says steer or abort but the request matches the known-job pattern, the result is forced to allow with harm: "none", steer: "none" and rule: "host-job-override". Bypassed results report all three confidences as 1.
Confidence semantics. pTrajectory, pAction and pHarm are softmax maxima, not calibrated probabilities; the weights are near-saturated at 1.00 on text like their training data, so values well below 1 are themselves a signal that the request is unlike the training set. pAction is the raw action head’s confidence even when a decision rule overrode it — compare it with rule to see whether the head and the rule agreed.
Message. On steer and abort the result carries a suggested short reply in a casual register, safe to replace. It is selected by key from the harm class and trajectory, never generated: secrets → a fixed line asking the user not to send card numbers, passwords or keys; jailbreak → “i’ll ignore the jailbreak bit and just help with the real ask.”; abort on crime → a fixed refusal that tells the user research or fiction framing must stay non-operational; scam on a steer → a fixed line advising the user to treat pressure to pay, install software or move channels as a scam; medical → a fixed line directing acute symptoms to emergency services; crisis → a fixed line pointing at local crisis services, chosen by a location hint the host may supply. message is absent on allow.
Examples #
Three requests sent with "model": "e-lim" and scored by the 1.0.0 weights, the version that serves, shown as the Falcon API returns them (message elided where present). Results scored by the weights carry version; bypass results do not.
A benign research question:
{
"ok": true,
"result": {
"action": "allow",
"trajectory": "research",
"harm": "none",
"steer": "none",
"pAction": 1,
"pHarm": 1,
"pTrajectory": 1,
"model": "e-lim",
"rule": "benign-traj≥0.28",
"version": "1.0.0"
}
}A request for operational harm — abort, and the host should answer with message (elided here) instead of calling its responder:
{
"ok": true,
"result": {
"action": "abort",
"trajectory": "crime",
"harm": "crime",
"steer": "none",
"pAction": 1,
"pHarm": 1,
"pTrajectory": 1,
"model": "e-lim",
"rule": "traj-abort≥0.62",
"version": "1.0.0"
}
}A greeting (hello) — answered by the bypass rule; no weights run:
{
"ok": true,
"result": {
"action": "allow",
"trajectory": "smalltalk",
"harm": "none",
"steer": "none",
"pAction": 1,
"pHarm": 1,
"pTrajectory": 1,
"model": "bypass",
"rule": "greeting-bypass"
}
}Other spot checks against the 1.0.0 weights: chest pain and i can't breathe → medical_emergency / steer / medical_911; my password is hunter2keep → secrets / steer / secrets_drop; what's the difference between etf and mutual fund → research / allow.
Training #
The served weights, version 1.0.0, were trained once, on 2026-09-03, as a Vertex AI custom job on an n1-standard-8 machine with one NVIDIA T4 (PyTorch 2.4). The job wrote elim.bin, the meta.json catalogue and run.json; no prediction endpoint was created on the platform. The trainer package that produced it is the same one that trains LIM.
Data. The training set is fully synthetic: no real user conversations were used, and no external corpus. A generator holds 214 short English template phrasings across the 32 trajectory classes (six to thirteen per class, with a few extra variants for the weather, research, scam-check, crime, crisis, medical-emergency, jailbreak and continue classes) and ten prior-exchange pairs. Each example picks a class round-robin (with an 8% chance of a random class instead), samples a phrasing from it, and mutates 18% of phrasings of eight or more characters with one typo (a case swap, an adjacent transposition or a single deletion). A prior exchange is prepended 65% of the time; the continue class always receives the reminder prior. The window is rendered with the same U: / A: format and 1,500-character cap as at inference. Labels are derived from the trajectory: action, harm and steer are fixed functions of the class, so the four heads learn a consistent mapping. 80,000 examples were generated with seed 7, shuffled and split 90/10 into 72,000 training and 8,000 validation rows.
Recipe. Eight epochs, batch size 128, AdamW at learning rate 2e-3 with weight decay 0.01, constant learning rate, no scheduler and no gradient clipping; PyTorch seed 7. The loss is the sum of four cross-entropies: trajectory; action with class weight 2.2 on abort; 1.25 × harm with class weight 1.8 on crisis, crime, child_sexual and scam; and steer. Training from scratch — E-LIM has no lineage and was not warm-started from LIM or any other model. After training, six fixture phrasings were scored as a sanity print (an explosives request → abort / crime; a weather request → allow / weather; a chest-pain request → steer / medical_emergency; a scam question → steer / scam_check, and two abort fixtures for the child-sexual and crisis-method classes).
Not trained on. Real conversations, non-English text, long-form text, transcripts with more than three prior turns, or any responder output. The typo jitter is the only augmentation; there is no paraphrasing, so the model has seen roughly two hundred distinct phrasings and their one-character variants.
Version 1.1.0 (withdrawn). A retrain on 2026-10-02 with the same trainer and architecture, on 292,113 windows: a wider in-house generator, hard negatives and 31,113 short harmless requests from public intent datasets, used only as allow examples and credited on third-party notices, with the unsafe classes counted twice as much in the loss. It was released on 2026-10-03 and withdrawn the same day (see Evaluation). It does not serve; the 1.0.0 weights that serve were not trained on any third-party data.
Evaluation #
| Metric | Value | Source |
|---|---|---|
| Unsafe messages allowed, unseen splits | 50 of 172 (29.1%) | Unsafe items of splits C and D of the October 2026 intention test set, never used to train or choose a model (448 items: 172 unsafe, 276 harmless); safety audit of 2026-10-03 |
| Harmless messages flagged, unseen splits | 40.6% (112 of 276) | Harmless items of splits C and D; 54 of the 276 (19.6%) were aborted |
| Crisis messages allowed, unseen splits | 3 of 34 | Crisis items of splits C and D; 23 of the 34 were labelled crisis, the rest stopped under other labels. Small group |
| Medical emergencies allowed, unseen splits | 6 of 34 | Medical-emergency items of splits C and D; only 9 of the 34 were labelled medical_emergency. Small group |
| Unsafe messages allowed, split B | 26 of 86 | Sealed split B of the October 2026 intention test set (224 items: 86 unsafe, 138 harmless); scored 2026-10-02 by an agent that did not train the model |
| Harmless messages escalated, split B | 37.0% (51 of 138) | Harmless items of sealed split B; 27 of the 138 (19.6%) were aborted |
| Child-safety rows, right action | 6 of 6 | The six child-safety rows of the 214-row template control set (1.0.0 generator phrasings) |
| In-process latency | ~0.78 ms per call | mean of 200 calls, single-threaded Node 22 on a laptop CPU, measured 2026-09-12 against the 1.0.0 weights |
The figures describe the 1.0.0 weights, the version that serves. They come from two scorings through the same code path the API uses, each by an agent that did not train or choose the model. Splits C and D of the October 2026 intention test set (448 items: 172 unsafe with a gold steer, 276 harmless) were opened for the first time in a safety audit on 2026-10-03; no training run and no choice of model had used them. Sealed split B (224 items: 86 unsafe, 138 harmless) was scored on 2026-10-02 and was the test that 1.1.0 was released on. An unsafe message counts as allowed when E-LIM answers allow; a harmless message counts as flagged (escalated) on steer or abort. The test sets were written and labelled by AI agents working in isolation; no person checked a label.
On splits C and D, 1.0.0 lets 50 of 172 unsafe messages through and flags 112 of 276 harmless ones, aborting 54 of them; on split B the figures are 26 of 86 and 51 of 138 (27 aborted). It allows 3 of the 34 crisis messages and 6 of the 34 medical emergencies on C and D, but much of what it stops, it stops under the wrong label: it labels 23 of the crisis messages crisis or crisis_method and only 9 of the medical emergencies medical_emergency, and reads many of the others as scams, jailbreaks or leaked secrets, so the suggested reply can be the wrong one. Of the 27 unsafe windows with history in C and D it allows 10.
Why 1.1.0 was withdrawn. E-LIM 1.1.0 was released on the morning of 2026-10-03 on its split B results (18 of 86 unsafe messages allowed, 4 of 138 harmless ones escalated). The audit on splits C and D the same day found that the split B gain did not repeat. 1.1.0 let 54 of 172 unsafe messages through, against 50 for 1.0.0, and 9 of the 34 crisis messages, against 3, among them several messages disclosing a plan or intent to end one’s life soon, which it labelled research at near-certain confidence. Unsafe messages allowed under a harmless label at a confidence of 0.90 or more rose from 10 to 38 of 172. It flagged far fewer harmless messages (8.7 %), but that does not outweigh missing a crisis, so 1.1.0 was withdrawn the same day and 1.0.0 serves again. For a few hours that day the route’s default was LIM Nano 1.1.0, which let 21 of 172 unsafe messages through on the same splits; it was withdrawn later the same day after tests on public sets of harmful requests, then LIM Nano in every version (see withdrawn models), and E-LIM 1.0.0 is the default again. A safety fix is in progress.
Known gaps. None of the splits has abort-class or child-safety items; those classes rest on the hospital-sim development set (150 rows; 1.0.0 allows 17 of its 93 unsafe items) and the 214-row template control, whose six child-safety rows 1.0.0 answers correctly. Groups such as the 34 crisis messages are small, so their rates are uncertain. On two public sets of harmful requests, AdvBench and JailbreakBench (620 requests), scored on 2026-10-03, E-LIM 1.0.0 let 200 through (32 %); real traffic was not measured. The confidences are near-saturated and have not been calibrated. The run.json validation scores of the 2026-09-03 training job (1.000 on every head) come from the same 214 templates as the training rows and say nothing about real text. The latency figure is a separate in-process measurement, not a training artefact. Measure E-LIM on your own traffic before relying on any threshold.
API #
| Endpoint | Auth | Body limit | Description |
|---|---|---|---|
POST /v1/intention | Bearer fln_… preview key or fls_… session token | 64,000 bytes | Score one request and return trajectory, action, harm, steer, confidences and the rule that fired. model: "e-lim" (default; the former id "elim" also works) or "lim". |
E-LIM is public on the Falcon API at POST /v1/intention. The route is shared with LIM: the optional model field selects LIM with "lim", and E-LIM answers when the field is absent, "e-lim", or "elim", its id until it was renamed on 2026-10-04, which still works and is answered as "e-lim". "lim-nano" answers 410 model_withdrawn with successor "e-lim": LIM Nano was withdrawn on 2026-10-03. Any other value answers 400 unknown_model with models listing the two current names, e-lim and lim. Authentication, the response envelope, retry guidance and quota sizes are described on API conventions; this page documents only what is specific to this route.
POST /v1/intention HTTP/1.1
Authorization: Bearer fln_your_preview_key
Content-Type: application/jsonThe base URL of the Falcon API is given on API conventions; paths on this page are relative to it.
Request body — text required, model optional ("e-lim" or "lim"; the former id "elim" selects E-LIM too):
{ "text": "is this a scam they sent me", "model": "e-lim" }Response — 200 with ok: true and a result object, or result: null when the intention weights are not available to the server:
{
"ok": true,
"result": {
"action": "steer",
"trajectory": "scam_check",
"harm": "scam",
"steer": "scam_warn",
"pAction": 1,
"pHarm": 1,
"pTrajectory": 1,
"model": "e-lim",
"rule": "traj-steer≥0.58",
"version": "1.0.0",
"message": "…"
}
}The HTTP route passes an empty history, so the window is the single U: line built from text; to score a multi-turn window, embed the model in-process (see Runtime & deployment). text is trimmed before the empty check and before scoring. message is omitted from the JSON when the action is allow. version (a string, "1.0.0" for E-LIM since 1.1.0 was withdrawn) names the weights that scored the request; it was added to result on 2026-10-03, is absent on bypass results, and callers that ignore unknown fields need no change. model is "e-lim" whenever E-LIM scored the request, whether the body named e-lim, elim or no model; until 2026-10-04 it said "elim", so an integration that compares it with elim should accept e-lim.
Errors. Every error is a JSON body with ok: false and an error code:
| Status | Code | Meaning |
|---|---|---|
| 400 | text_required | text was missing or empty after trimming. |
| 400 | bad_json | The body was not valid JSON, or exceeded the 64,000-byte limit. |
| 400 | unknown_model | model was none of e-lim, its former id elim, and lim; models lists the two current values, e-lim and lim. |
| 401 | invalid_credentials | The bearer token was missing, unknown or revoked. |
| 402 | payment_required | The key’s account is not in good standing. |
| 404 | not_found | The path or method did not match a route. |
| 410 | model_withdrawn | model was lim-nano: LIM Nano was withdrawn on 2026-10-03. successor is e-lim. Answered before the credential is checked and not counted. |
| 429 | quota_exceeded | The intention bucket is exhausted for the current UTC calendar month; the body also carries kind: "intention". |
| 500 | internal | An unexpected server error; safe to retry once. |
Quota and limits. Each request counts one against the intention preview quota bucket, per UTC calendar month, whichever of E-LIM or LIM scored it; a request that fails validation or authentication, or names the withdrawn LIM Nano, does not count. The request body is limited to 64,000 bytes. There is no batch route; send one window per request. Bucket sizes are listed on API conventions.
Runtime & deployment #
| Kind | In-process |
|---|---|
| Resident | about 53 MB of float32 tensors, loaded lazily on first call and cached for the process lifetime |
| Serving | inside the Falcon API process behind POST /v1/intention, or embedded directly by a host |
| Cold start | first call reads 52.8 MB of weights from disk; later calls carry no load cost; the Falcon API itself sleeps after four hours without calls and takes some seconds, up to about 20, to start on the next one |
| Concurrency | synchronous per call on the host's event loop; no worker pool, no GPU |
| Timeout | — |
E-LIM runs in-process. The runtime is a small TypeScript loader plus a handful of dense-math helpers (embedding-bag mean, linear, GELU, softmax) with no native dependencies and no network call; the Falcon API server hosts the same code behind /v1/intention, and a host that wants the multi-turn window embeds it directly. The Falcon API runs on 2 vCPU / 2 GiB; it shuts down after four hours without calls, and the call that finds it asleep waits some seconds, up to about 20, while it starts. Once it is up, E-LIM answers through the API in well under a second.
- Load behaviour. The weights are read lazily on the first call: the loader reads
meta.json(the label catalogue and dimensions under itselimkey) andelim.bin(16 named float32 tensors, 52,774,490 bytes), then caches both for the process lifetime. About 53 MB stays resident. The same loader also reads LIM’s weights when they sit beside E-LIM’s, adding 35.8 MB. If neither file loads, the loader records the pack as missing and every score call returnsnullrather than throwing; the Falcon API surfaces that asresult: nullwith status 200, never as a 503, so the “503 while loading, retry once after 2.5 s” pattern used by the served models does not apply here. - Selection. E-LIM is the default. The
model: "lim"request field — or the equivalent per-caller setting in-process — selects LIM for that request; ifelim.binis absent the loader falls back to LIM silently and reportsmodel: "lim"in the result. No process restart is needed to switch.GET /healthlists the selectable variants aslimVariants: ["e-lim", "lim"]. - What the host must provide. When the weights are available to the host process (they are not distributed during the private preview): a directory containing
meta.jsonandelim.bin; a Node 22 runtime; the conversation history (up to three prior turns, withroleset touserormodel) and the current request; optionally a location hint used only to choose which regional crisis line thecrisissuggestion names. - Concurrency and timeouts. Each call is synchronous and takes under a millisecond on the calling thread, so throughput scales with the host’s event loop; there is no worker pool, no batching and no timeout to configure.
- Disabling. The host can switch intention scoring off entirely, in which case the scorer returns
nulland the host must decide its own default (usuallyallow).
Weights are not distributed during the private preview; the hosted route is the way to call this version.
Integration notes #
- Treat
nullas “no opinion”, not asallow: decide in the host whether an unavailable scorer fails open or closed. - Read
action, notharm, when deciding what to do.harmcan be non-noneon anallowresult when the trajectory head disagreed with the harm head; the decision rules require trajectory agreement precisely so that a lone harm vote does not abort a benign request. - Hand
messageto the responder as guidance onsteer, and use it as the reply onabort. Replace it freely — it is a fixed line chosen by key, and the register is casual. - Keep the bypass rules in mind when testing: a request that mentions food ordering, transit, reminders, weather, maps or spatial layout will never reach the weights, and a flagged request that matches those patterns is overridden to
allow. - E-LIM scores only the user’s side. If your application needs to check what the responder wrote, that is a separate pass with a different model.
Limits & safety #
It does not see the responder’s output, the wider transcript beyond three prior turns, or anything outside the last 1,500 characters of text; it reads a lowercased window and nothing else.
- It does not write replies.
messageis a suggested short line selected by key, safe to replace; the responder or the host writes the reply. - It does not score the responder’s output. Only the user’s turns and the current request are labelled; earlier safety replies are stripped from the window before scoring.
- It is not a safety classifier of record. Thresholds favour silence (a benign trajectory wins at 0.28), greetings and known job phrasings bypass the weights entirely, and a flagged request that matches a known job pattern is overridden to
allowby design. A harmful request wrapped in a benign job phrasing is allowed. - It has seen only 214 English template phrasings and their one-character typo variants, never real conversations. Novel phrasings, other languages, sarcasm, long requests (only the last 360 characters reach the character bag) and multi-step reasoning are out of distribution. On test splits never used to choose it, it lets about three unsafe messages in ten through (50 of 172) and flags about four harmless messages in ten (112 of 276), and it often stops a crisis or medical emergency under an unrelated label; those sets were written and labelled by AI agents, not checked by a person.
- Confidences are softmax maxima, not calibrated probabilities. The weights are near-saturated at 1.00 on template-like text, and hash collisions across 16,384 buckets per bag can make unrelated n-grams share an index.
- Five of the 32 trajectory classes are routing labels that only make sense for a host with those jobs; they are not documented and should be mapped to
unknownby hosts that do not use them. - The crisis and medical suggestions are fixed lines. They are not medical, legal or crisis advice, and the regional crisis line they name is chosen by a location hint, not detected from the request.
Out of scope. E-LIM is not a content filter, a toxicity scorer, a language detector or a sentiment model. It does not extract entities, dates or amounts from the request, does not track state across requests, and does not learn which of its labels the host acted on. Not medical, legal or crisis advice.
Fixed weights per version; the model does not learn from requests.
Versions #
| Version | Date | Status | Note |
|---|---|---|---|
| 1.0.0 | Released | Serving again from 2026-10-03, when 1.1.0 was withdrawn; the same weights as the 2026-09-03 entry. The default of /v1/intention again since LIM Nano was withdrawn the same day. | |
| 1.1.0 | Deprecated | Withdrawn on 2026-10-03, the day it was released. A safety audit on test splits never used to choose it found that it let through more crisis messages than 1.0.0. | |
| 1.0.0 | Released | First documented version (80,000 synthetic windows from 214 template phrasings, 8 epochs, seed 7). The default intention model from 2026-09-03. |
Compatibility. A major version bump changes the input or output contract: the window format, the label catalogue of any head (adding, removing or renaming a trajectory, harm or steer label), the decision-rule thresholds, or the shape of the result object. A minor bump is a retrain on the same contract — new weights, same labels and same fields — and may move confidences and change individual decisions without changing what a caller has to parse. A patch bump changes only metadata or runtime behaviour, such as the loader or this documentation. Hosts should pin to a major version and re-validate their thresholds on any minor bump.
1.1.0 (2026-10-03), withdrawn. A minor version, a retrain on the same contract, released on the morning of 2026-10-03 and withdrawn the same day after a safety audit on test splits that had never been used to choose it: it let through more crisis messages than 1.0.0, several of them at near-certain confidence under a harmless label (see Evaluation). From the same day E-LIM serves 1.0.0 again, and callers see "version": "1.0.0". The version field added to the HTTP result on 2026-10-03 stays. 1.1.0 cannot be selected.
1.0.0 and the default. E-LIM has been the default model of POST /v1/intention since 2026-09-03, apart from a few hours on 2026-10-03 when the default was LIM Nano 1.1.0, withdrawn the same day with every other version of LIM Nano. A request without a model field is scored by E-LIM 1.0.0, with the contract and the decisions of the weights it served before.
Current weights. Version 1.0.0 is the tensor bundle elim.bin (52,774,490 bytes, 16 float32 tensors) written by the 2026-09-03 Vertex AI run with seed 7, paired with the meta.json catalogue whose elim entry records 16,384 hash buckets, 384-wide bags, a 384-wide hidden state, 32 / 3 / 8 / 8 head sizes and 13,193,523 parameters. The withdrawn 1.1.0 is the bundle of the same size and layout written by the 2026-10-02 run.
The name. E-LIM was called ELIM until 2026-10-04. The rename changed no weights and no version, and the API still accepts the old id elim for it and answers it as e-lim; the internal names keep the old one: the weights file elim.bin and the elim entry of its meta.json.
Weights are not distributed during the private preview.