Skip to content

@memhtml/llm

@memhtml/llm: the Bedrock embeddings and structured-output lanes every memhtml model call goes through, over one client type that is satisfied by Bedrock directly or by an HTTP LLM proxy.

The embeddings lane is Cohere Embed v4 on InvokeModel. Embeddings is the service and EmbeddingsShape has two entry points: embed sends documents with input_type search_document, and embedQuery sends one retrieval query with input_type search_query, because Cohere embeds the two into different regions of the same space and a corpus indexed one way must be queried the other way. chunkTexts slices a document list into batches of EMBED_BATCH_LIMIT (96) texts, makeEmbeddings runs up to EMBED_CONCURRENCY (6) batches at once and flattens the vectors back in input order, and buildEmbedBody names output_dimension as EMBED_DIM (1024) because the model returns 1536 floats when the field is absent. readEmbeddings refuses a response whose vector count or vector width disagrees with the request, since a short answer would shift every later vector onto the wrong chunk. EMBED_WATERMARK joins the model id and the dimension into cohere.embed-v4:0@1024, the value index_state.embed_model stores, and @memhtml/index, @memhtml/eval and apps/cli import it from here rather than concatenating their own.

The generation lane is ModelClient, whose ModelClientShape has generate for prose (returning a Generation with text, token counts and latency) and generateObject for one schema-shaped answer described by a StructuredRequest. MODELS is the table of five ModelKey values, three Anthropic (sonnet-5, opus-5, fable-5) and two OpenAI (gpt-5.6-sol, gpt-5.6-terra), each carrying a global. inference-profile id and a provider, and modelByKey resolves a key to its row. Effort is the four-value reasoning dial (low, medium, high, xhigh); the Anthropic body sends it as output_config.effort and the OpenAI body as reasoning_effort. thinkingFor returns {type: "adaptive"} for opus-5 and fable-5 and null for sonnet-5, which rejects a thinking key instead of ignoring it. buildInvokeBody writes the body in the model’s dialect: the Anthropic Messages body with anthropic_version set to ANTHROPIC_VERSION and, for a structured call, a forced tool named STRUCTURED_TOOL_NAME (emit); or the OpenAI chat-completions body with response_format set to a strict json_schema under the same name. clampTokens bounds the output budget between 1 and MAX_TOKENS_CEILING (128,000), defaulting to MAX_TOKENS_DEFAULT (64,000), because Bedrock rejects rather than clamps. On the read side both dialects are folded into one InvokeResponseBody, incompleteReason fails a response whose stop_reason is in INCOMPLETE_STOP_REASONS (max_tokens or refusal) before any content is read, readText joins the text blocks and drops thinking blocks, and readToolInput finds the emit block by name rather than by position. toInputSchema derives the tool’s JSON Schema from an effect Schema, folding hoisted definitions back under $defs, and decodeToolInput decodes the tool payload with excess properties rejected, allowing one repair for a top-level container field that arrived double-encoded as a JSON string. wrapAsData wraps memory or rollout text in a labeled data block and neutralizes any closing tag inside it, so instruction-shaped corpus prose cannot end the block early.

Both lanes are written against InvokeClient, a structural type with one send method resolving to an InvokeResult that a BedrockRuntimeClient and a recording fake both satisfy, and invokeJson is the one round trip: send the body, decode the JSON payload. LlmConfig reads MEMHTML_AWS_REGION (default us-east-1) and the four proxy variables; makeBedrockClient builds the direct client with maxAttempts: 10 adaptive retry and REQUEST_HANDLER_OPTIONS, whose requestTimeout of 300,000 ms turns a hung socket into a retryable failure. proxyConfigFromEnv is the plain-function reader of the same variables: MEMHTML_LLM_BASE_URL names the proxy origin and is null when absent or blank, which means Bedrock directly; MEMHTML_LLM_API_KEY is an optional bearer token; MEMHTML_LLM_MODEL_PREFIX is the prefix put in front of every Bedrock id, DEFAULT_PROXY_MODEL_PREFIX (bedrock/) when blank and no prefix when it is the word PROXY_MODEL_PREFIX_NONE (none); and MEMHTML_LLM_MODEL_MAP is comma-separated from=to pairs that parseProxyModelMap reads and proxyModelId applies ahead of the prefix. A set-but-malformed origin or map entry throws at construction with the variable named, so a typo cannot silently route traffic back to Bedrock.

invokeClientFor is the one construction site: makeProxyClient when the configuration names a proxy, makeBedrockClient otherwise, and both ModelClientLive and EmbeddingsLive call it so the two lanes cannot disagree about where a run’s traffic goes. makeProxyClient satisfies InvokeClient over fetch, typed as ProxyFetch returning a ProxyFetchResponse so a test can hand it a recorder: proxyRouteFor picks the route by model id from PROXY_ROUTE_PATHS (/v1/messages for Anthropic models, /v1/chat/completions for OpenAI models, /v1/embeddings for the embedder), toProxyRequest drops anthropic_version and adds model on the Messages route, adds model on the completions route, and rewrites Cohere’s texts and output_dimension into OpenAI’s input and dimensions on the embeddings route, and fromProxyResponse folds an OpenAI embeddings answer back into the Cohere shape readEmbeddings reads. A non-2xx answer becomes a ProxyHttpError carrying the status and the body’s first 200 characters, and isRetryableProxyFailure retries 408, 429 and 5xx under exponential backoff with jitter bounded to about two minutes. apps/consolidator keeps its own dependency-free copy of the proxy-config.ts parsers because its agent file cannot import a workspace package.

Two error types from @memhtml/contracts divide the failures. ModelUnavailable means no usable answer arrived: a transport rejection or a payload that fails to parse (reduced by modelFailure to the model id and name: message), an incomplete stop_reason, an embedding count or width mismatch, or a prose call that returned no text. LlmContractViolation means an answer arrived but broke the structured contract: no emit tool call, or a payload that fails the strict decode, with the raw payload capped at MAX_RAW (800) characters in the reason.

The package publishes one import path, @memhtml/llm, which re-exports the nine modules below, so the reference on this page is the whole exported surface. One module-level export is not re-exported: wire.ts’s normalizeOpenAiResponse, which model-client.ts uses internally to fold an OpenAI payload into InvokeResponseBody.

Module What it holds
client.ts InvokeClient, InvokeResult, invokeJson, modelFailure, REQUEST_HANDLER_OPTIONS, makeBedrockClient, and the LlmConfig reader with its LlmConfigShape.
constants.ts The embed wire constants (EMBED_MODEL_ID, EMBED_DIM, EMBED_BATCH_LIMIT, EMBED_CONCURRENCY, EMBED_WATERMARK) and the generation ones (STRUCTURED_TOOL_NAME, MAX_TOKENS_DEFAULT, MAX_TOKENS_CEILING, ANTHROPIC_VERSION).
embeddings.ts Embeddings, EmbeddingsShape, EmbeddingsLive, makeEmbeddings, and the pure pieces chunkTexts, buildEmbedBody and readEmbeddings.
model-client.ts ModelClient, ModelClientShape, ModelClientLive, makeModelClient, the Generation and StructuredRequest shapes, and wrapAsData.
models.ts MODELS, ModelInfo, ModelKey, Provider, modelByKey, Effort and thinkingFor.
proxy-config.ts The variable names (PROXY_BASE_URL_VAR, PROXY_API_KEY_VAR, PROXY_MODEL_PREFIX_VAR, PROXY_MODEL_MAP_VAR), DEFAULT_PROXY_MODEL_PREFIX, PROXY_MODEL_PREFIX_NONE, ProxyConfig, and the parsers normalizeProxyBaseUrl, parseProxyModelMap, proxyModelPrefix, proxyConfigFromEnv and proxyModelId.
proxy.ts ProxyRoute, PROXY_ROUTE_PATHS, proxyRouteFor, ProxyRequest, toProxyRequest, fromProxyResponse, ProxyFetch, ProxyFetchResponse, ProxyHttpError, isRetryableProxyFailure, ProxyClientOptions, makeProxyClient and invokeClientFor.
structured.ts toInputSchema, decodeToolInput and MAX_RAW.
wire.ts GenerateOptions, clampTokens, JsonSchemaObject, StructuredTool, buildInvokeBody, INCOMPLETE_STOP_REASONS, ContentBlock, InvokeResponseBody, asResponseBody, incompleteReason, readText and readToolInput.

A non-2xx answer, carrying the status and the body’s first 200 characters. The body is kept because a proxy reports routing and quota failures as structured JSON an operator needs verbatim — {"error":{"code":"model_not_found"}} names the fix, where a bare status does not.

  • Error
new ProxyHttpError(status, body): ProxyHttpError;

number

string

ProxyHttpError

Error.constructor
readonly name: "ProxyHttpError" = "ProxyHttpError";
Error.name
readonly status: number;
readonly optional input?: unknown;
readonly optional name?: string;
readonly optional text?: string;
readonly optional type?: string;

Cohere Embed v4 on bedrock-runtime InvokeModel.

Two entry points, and the difference changes the result. Cohere embeds documents and queries into deliberately different regions of the same space, so a corpus indexed with search_document must be queried with search_query for the cosine to mean what the retrieval arm assumes. Reusing one input_type for both degrades every vector hit without failing anything.

readonly embed: (texts) => Effect<readonly Float32Array<ArrayBufferLike>[], ModelUnavailable>;

Embed documents for storage, in order, chunking requests at Cohere’s ceiling.

readonly string[]

Effect<readonly Float32Array<ArrayBufferLike>[], ModelUnavailable>

readonly embedQuery: (text) => Effect<Float32Array<ArrayBufferLike>, ModelUnavailable>;

Embed one retrieval query. Same space as embed, different input_type.

string

Effect<Float32Array<ArrayBufferLike>, ModelUnavailable>


The invoke_model bodies, one dialect per provider, and the read-side that folds both dialects into ONE response shape.

The Anthropic lane is the native Messages body. The effort and thinking rules are per-model and exact, and Converse has no field for either, so Converse is not used.

The OpenAI lane is the chat-completions body, and its structured mechanism is response_format: {type: "json_schema", strict: true} — constrained decoding, so the bytes cannot leave the schema. Its response is normalized into InvokeResponseBody right here at the wire, with the schema-constrained answer presented as a tool_use block named emit: every consumer from readToolInput through decodeToolInput then has exactly one shape to read, and the decode stays the single gate for both providers.

readonly optional cacheSystem?: boolean;

Mark the system prompt as a cache breakpoint.

The batched phases send one system prompt and one tool schema across every batch of a night, with only the member list changing per call, so the prefix is the same bytes tens of times in a row. With this set, system goes out as a content-block array carrying cache_control: {type: "ephemeral"} instead of a plain string, which is how the Messages API names a prefix to cache. A plain string carries no place to put the marker, so the shape has to change and not only gain a field.

readonly effort: "low" | "medium" | "high" | "xhigh";
readonly optional maxTokens?: number;
readonly optional openaiPromptCache?: OpenAiPromptCache;

How the OpenAI dialect asks Bedrock to treat prompt caching. See OpenAiPromptCache. Absent means DEFAULT_OPENAI_PROMPT_CACHE. The Anthropic dialect ignores it: that lane names its cache prefix with cacheSystem and has no automatic breakpoint to turn off.

readonly optional system?: string;

The two ways a sleep phase reaches a model: prose, and one forced-tool object.

Per-item failure isolation is NOT here. A phase that iterates candidates decides for itself whether one bad response skips an item or fails the phase, and it does that with Effect.result over each call. This service’s contract is narrower and total. One call goes in, and either a value that honors its type or a typed failure comes out.

readonly inputTokens: number | null;
readonly latencyMs: number;
readonly outputTokens: number | null;
readonly text: string;

The one Bedrock call this package makes, named as a structural type instead of the SDK class. BedrockRuntimeClient satisfies it, and so does a fake that records the request body and returns a canned payload. That fake lets every wire assertion (batch boundaries, output_dimension, the thinking key, a truncated stop_reason) run with no network and no credential. The response is narrowed to body because that is the only field either lane reads.

readonly send: (command, options) => Promise<InvokeResult>;

InvokeModelCommand

AbortSignal

Promise<InvokeResult>


readonly optional content?: readonly ContentBlock[];
readonly optional stop_reason?: string | null;
readonly optional usage?: object;
readonly optional input_tokens?: number;
readonly optional output_tokens?: number;

What InvokeClient.send resolves to: the SDK’s InvokeModel output narrowed to the one field either lane reads. BedrockRuntimeClient’s own output type satisfies it.

readonly optional body?: Uint8Array<ArrayBufferLike>;

The raw response payload, absent when the SDK returned no body.


A JSON Schema object for the forced tool’s input_schema. Open-keyed because JSON Schema is, and because the derivation in structured.ts emits keys this module has no reason to enumerate.

[key: string]: unknown

readonly openaiPromptCache: OpenAiPromptCache;
readonly proxy: ProxyConfig | null;
readonly region: string;

Per-deployment wire defaults a caller’s GenerateOptions does not name. The sleep phases say what a call IS (system, effort, budget); the environment says how the transport should carry it, and this is where the two meet. A field set on the call wins over the default.

readonly optional openaiPromptCache?: OpenAiPromptCache;

See OpenAiPromptCache; LlmConfig reads it from MEMHTML_OPENAI_PROMPT_CACHE.


readonly generate: (modelKey, prompt, options) => Effect<Generation, ModelUnavailable>;

"sonnet-5" | "opus-5" | "fable-5" | "gpt-5.6-sol" | "gpt-5.6-terra"

string

GenerateOptions

Effect<Generation, ModelUnavailable>

readonly generateObject: <A, I>(request) => Effect<A, ModelUnavailable | LlmContractViolation>;

A

I

StructuredRequest<A, I>

Effect<A, ModelUnavailable | LlmContractViolation>


readonly key: "sonnet-5" | "opus-5" | "fable-5" | "gpt-5.6-sol" | "gpt-5.6-terra";
readonly label: string;
readonly modelId: string;
readonly provider: Provider;

readonly optional fetch?: ProxyFetch;
readonly optional schedule?: Schedule<unknown, unknown, never, never>;

The retry schedule; a test substitutes one with no delay.


readonly apiKey: string | null;

null when the proxy takes no credential.

readonly baseUrl: string;

The origin with no trailing slash; routes are appended to it.

readonly modelMap: ReadonlyMap<string, string>;

memhtml’s model id → the exact id the proxy wants, sent without the prefix.

readonly modelPrefix: string;

Prepended to every Bedrock id the map does not name. May be empty.


The part of a fetch Response the proxy client reads; the platform Response satisfies it.

readonly ok: boolean;

true for a 2xx status, as the platform fetch reports it.

readonly status: number;

The HTTP status code the proxy answered with.

readonly text: () => Promise<string>;

Reads the whole response body as text, once.

Promise<string>


readonly body: Record<string, unknown>;
readonly path: string;
readonly route: ProxyRoute;

A

I

readonly optional cacheSystem?: boolean;

Cache the system prompt as a prefix across calls. See GenerateOptions.cacheSystem.

Set by callers that repeat one system prompt over many calls, which is every batched sleep phase. The value only reshapes the request body; the response and the decode are unchanged, so setting it can change what a call costs and cannot change what it returns.

readonly effort: "low" | "medium" | "high" | "xhigh";
readonly optional inputSchema?: JsonSchemaObject;

A hand-written input_schema, when the derived one is wrong for the case. The schema still decodes the response, so an override widens what the model is asked for without widening what is accepted.

readonly optional maxTokens?: number;
readonly modelKey: "sonnet-5" | "opus-5" | "fable-5" | "gpt-5.6-sol" | "gpt-5.6-terra";
readonly prompt: string;
readonly schema: Codec<A, I>;
readonly optional system?: string;
readonly optional toolDescription?: string;

The tool description. Set it, since it is the model’s only prose about the shape.


readonly optional description?: string;
readonly inputSchema: JsonSchemaObject;
type Effort = typeof Effort.Type;

Reasoning effort. The Anthropic lane passes it as output_config.effort, the OpenAI lane as reasoning_effort; both accept all four values (sol probed live 2026-08-22).


type ModelKey = typeof ModelKey.Type;

type OpenAiPromptCache = "off" | "implicit";

The two things the OpenAI dialect can say about prompt caching.

off is the default because of what the sleep prompts are. Bedrock’s implicit mode for GPT-5.6 places an automatic breakpoint on the LATEST message, so with a 1,024-token minimum every call whose whole prompt clears the minimum writes that whole prompt to the cache at 1.25x the input rate; and the eight sleep system prompts are 189 to 496 tokens, so the only prefix that could be reused never reaches the minimum and the only prefix that does reach it — system plus a unique batch — is never seen twice. Measured 2026-09-11 on the access log: every OpenAI-lane call reported cache_write_tokens about equal to its input and cached_tokens: 0, a pure premium with no read to pay it back. implicit is for a request whose prefix WILL repeat and clear the minimum, and for an OpenAI-compatible endpoint that rejects the Bedrock-only field.

Probed live 2026-09-11 against bedrock-runtime.us-east-1.amazonaws.com/openai/v1/chat/completions with global.openai.gpt-5.6-sol: the field is accepted on Chat Completions (200) and honored (prompt_tokens: 3424, cache_write_tokens: 0 on a prompt that wrote 2,641 tokens without it).


type Provider = "anthropic" | "openai";

Which wire dialect a model speaks. Selected per model, never per call.


type ProxyFetch = (url, init) => Promise<ProxyFetchResponse>;

The minimal fetch this client needs, so a test can hand it a recorder.

string

string

Record<string, string>

"POST"

AbortSignal

Promise<ProxyFetchResponse>


type ProxyRoute = "messages" | "completions" | "embeddings";

Which of the proxy’s routes a model id is served on.

const ANTHROPIC_VERSION: "bedrock-2023-05-31" = "bedrock-2023-05-31";

The only valid anthropic_version. Not a model date.


const DEFAULT_OPENAI_PROMPT_CACHE: OpenAiPromptCache = "off";

const DEFAULT_PROXY_MODEL_PREFIX: "bedrock/" = "bedrock/";

LiteLLM’s provider prefix for Bedrock: bedrock/global.anthropic.claude-opus-5.


const Effort: Literals<readonly ["low", "medium", "high", "xhigh"]>;

Reasoning effort. The Anthropic lane passes it as output_config.effort, the OpenAI lane as reasoning_effort; both accept all four values (sol probed live 2026-08-22).


const EMBED_BATCH_LIMIT: 96 = 96;

Cohere’s per-request text ceiling. Batches larger than this are rejected.


const EMBED_CONCURRENCY: 6 = 6;

How many embed batches are in flight at once.

A whole-store pass is ceil(chunks / EMBED_BATCH_LIMIT) requests, about 105 for a 10k-chunk corpus. Issuing them one after another makes index rebuild --embed a serial chain of network round trips, which is the slowest thing this system does.

The concurrency is bounded, and the bound protects a shared quota rather than local resources. Bedrock rate-limits per account, so every caller in the account draws on one tokens-per-minute budget, and an unbounded fan-out would spend that budget in one burst and throttle every other consumer. Throttles that do occur are absorbed below Effect by the SDK’s adaptive retry (maxAttempts: 10), which backs off per request. A slightly-too-high bound therefore costs latency instead of failing the run.


const EMBED_DIM: 1024 = 1024;

const EMBED_MODEL_ID: "cohere.embed-v4:0" = "cohere.embed-v4:0";

Bedrock wire constants. Cohere Embed v4 returns 1536 floats when output_dimension is absent and exactly 1024 when it is named (probed live 2026-08-02), so the InvokeModel body names it. If the default changed without notice, every stored vector would be invalid against a schema that says 1024.


const EMBED_WATERMARK: "cohere.embed-v4:0@1024";

The watermark value index_state.embed_model stores. Both axes in one string, because a model id alone does not identify a vector space: the same id at another output_dimension produces vectors that are silently incomparable with the stored ones.


const Embeddings: Service<EmbeddingsShape, EmbeddingsShape>;

const EmbeddingsLive: Layer<EmbeddingsShape, ConfigError, never>;

The transport is whichever invokeClientFor selects: Bedrock directly (maxAttempts: 10 with adaptive retry) or the LLM proxy MEMHTML_LLM_BASE_URL names (jittered exponential backoff on throttles and upstream failures). The embed lane is why the retries matter: it issues hundreds of calls per index run, so a throttle that failed the run instead of backing off would make a full rebuild unreliable at exactly the corpus size where the rebuild matters. Both transports bound a hung socket (REQUEST_HANDLER_OPTIONS), so one dead connection inside a batch fan-out fails that batch instead of stalling the whole rebuild.

Through the proxy, output_dimension travels as the OpenAI dimensions field. A proxy that drops it returns Cohere’s 1536-wide default, which readEmbeddings refuses by width — a typed failure and a degraded search, never a vector of the wrong shape in the index.


const INCOMPLETE_STOP_REASONS: ReadonlySet<string>;

stop_reason values that mean the content is not a complete answer. Both become typed failures. A response cut off at max_tokens may never have reached the point that made it a judgment, and a refusal carries no judgment at all. Reading either as a finished result would be a silent data-quality bug, so neither is coerced into one.


const LlmConfig: Config<{
openaiPromptCache: OpenAiPromptCache;
proxy: ProxyConfig | null;
region: string;
}>;

Where every lane sends its calls.

region is what the direct path resolves against. us-east-1 is the default because it is where both cohere.embed-v4:0 and the global.anthropic.* inference profiles are reachable.

Auth for the direct path is deliberately absent: the SDK’s default chain picks up AWS_BEARER_TOKEN_BEDROCK from the environment when it is set, and falls back to the standard credential chain (instance role, profile, environment keys) otherwise. Naming a profile or a key here would break both paths.

proxy is the LLM proxy every lane goes through instead, when MEMHTML_LLM_BASE_URL names one (proxy-config.ts has the three variables and the routes). null is the default and means Bedrock directly; the region is then unused but still read, so flipping a deployment between the two paths is one variable and not a reshuffle. A set-but-malformed value fails here, at construction, with the variable named — see normalizeProxyBaseUrl for why it does not fall back to the direct path.


const MAX_RAW: 800 = 800;

Cap on the raw payload carried on a violation, so a runaway response cannot bloat it.


const MAX_TOKENS_CEILING: 128000 = 128_000;

Every Claude 5 generation tops out here. Above the ceiling Bedrock raises a ValidationException rather than clamping, so the clamp lives on this side.


const MAX_TOKENS_DEFAULT: 64000 = 64_000;

The per-call output budget when a caller names none. A budget of 8192 has been observed truncating structured responses mid-object, and a truncated structured response is a contract violation rather than a partial result. max_tokens bounds thinking and answer together, so a tight budget is consumed sooner than the answer’s length alone suggests, and every sleep phase here runs with reasoning on.

64,000 rather than the 16,384 that stood here: the same number the consolidator settled on for its own per-call ceiling (apps/consolidator/src/output-budget.ts, issue #113), half of the 128,000 ceiling below, and generous against the largest structured answer any phase asks for. A budget is a ceiling and not a spend — the model stops when the answer is done — so the cost of a high default is only paid by the answers that needed it.


const ModelClient: Service<ModelClientShape, ModelClientShape>;

const ModelClientLive: Layer<ModelClientShape, ConfigError, never>;

The transport is whichever invokeClientFor selects from the configuration: Bedrock directly, or the LLM proxy MEMHTML_LLM_BASE_URL names. Both bound a hung socket (REQUEST_HANDLER_OPTIONS) as well as retrying throttles. The sleep phases run their model calls sequentially, so without the bound one dead connection would stall a whole night’s phase rather than failing it.


const ModelKey: Literals<readonly ["sonnet-5", "opus-5", "fable-5", "gpt-5.6-sol", "gpt-5.6-terra"]>;

const MODELS: ReadonlyArray<ModelInfo>;

Bedrock ids use the global. inference profiles, which makes them reachable from a single region without provisioning per-region throughput. The OpenAI models REQUIRE the profile: the bare openai.gpt-5.6-* ids reject on-demand invocation outright (probed live 2026-08-22).


const OPENAI_PROMPT_CACHE_OFF_FIELD: object;

The wire field off emits. One constant, so the body and the tests spell it once.

readonly mode: "explicit" = "explicit";

const OPENAI_PROMPT_CACHE_VALUES: ReadonlyArray<OpenAiPromptCache>;

const OPENAI_PROMPT_CACHE_VAR: "MEMHTML_OPENAI_PROMPT_CACHE" = "MEMHTML_OPENAI_PROMPT_CACHE";

The environment variable openAiPromptCacheFrom reads.


const PROXY_API_KEY_VAR: "MEMHTML_LLM_API_KEY" = "MEMHTML_LLM_API_KEY";

A bearer token the proxy requires, sent as Authorization: Bearer <key>. Optional.


const PROXY_BASE_URL_VAR: "MEMHTML_LLM_BASE_URL" = "MEMHTML_LLM_BASE_URL";

The proxy’s origin, e.g. http://127.0.0.1:4000. Absent or blank means Bedrock directly.


const PROXY_MODEL_MAP_VAR: "MEMHTML_LLM_MODEL_MAP" = "MEMHTML_LLM_MODEL_MAP";

from=to pairs, comma-separated, naming single models to the proxy by exact id: cohere.embed-v4:0=cohere-embed-v4. A mapped id is sent verbatim, with no prefix.


const PROXY_MODEL_PREFIX_NONE: "none" = "none";

The value of PROXY_MODEL_PREFIX_VAR that means “no prefix”. A WORD rather than the empty string, because effect/Config reads an empty environment value as absent (probed against the pinned release: Config.String fails on "" exactly as on a missing key, so withDefault fires for both). The sleep lanes read this variable through Config and the consolidator reads it from process.env directly, and “” would have meant the default on one path and no prefix on the other. Compared case-insensitively.


const PROXY_MODEL_PREFIX_VAR: "MEMHTML_LLM_MODEL_PREFIX" = "MEMHTML_LLM_MODEL_PREFIX";

The prefix put in front of every unmapped Bedrock id. Unset or blank means DEFAULT_PROXY_MODEL_PREFIX; the literal PROXY_MODEL_PREFIX_NONE sends bare ids.


const PROXY_ROUTE_PATHS: Readonly<Record<ProxyRoute, string>>;

const REQUEST_HANDLER_OPTIONS: object;

Request-handler options for every Bedrock client this package constructs.

The SDK’s default request timeout is 0 — no bound at all — so a socket that hangs after the request is written never errors, and the call holding it stalls its caller forever. A sleep phase runs its calls sequentially, so one hung socket stalls the whole night.

This client’s default handler is NodeHttp2Handler (pinned SDK 3.1111.0, resolved at dist-es/runtimeConfig.js), and a plain options object here is passed to that handler’s constructor, so every key must be one that handler reads: requestTimeout, sessionTimeout, disableConcurrentStreams, maxConcurrentStreams, and nodeHttp2ConnectOptions. There is no connect timeout among them — a connectionTimeout is accepted by the type and never read.

requestTimeout is the one bound, and it is per-STREAM: the handler arms it with clientHttp2Stream.setTimeout, so it fires after that long with no activity ON THE CALL and rejects with a TimeoutError. That name is what @smithy/core’s service-error-classification matches, so the rejection is retryable and maxAttempts: 10 re-attempts it. 300s of inactivity is far above any legitimate single-token gap — a high-effort structured call streams nothing until it answers, and the slowest observed answers are minutes, not five — so the bound turns only genuinely dead sockets into typed failures.

sessionTimeout is deliberately absent, because it does not bound only an idle session. Probed 2026-08-25 against a loopback h2 server: NodeHttp2Handler passes sessionTimeout as the connection manager’s connectionConfiguration.requestTimeout, which arms session.setTimeout(value, ensureDestroyed), and ClientHttp2SessionRef.destroy() tears the session down with no in-flight check. Waiting for a slow answer IS “no activity” on the session, so the timer fires mid-request: against a server answering at 400 ms, a 50 ms sessionTimeout rejected the call at 71 ms with a bare Error: Unexpected error: http2 request did not get a response carrying no code and no $metadata — a shape no retry predicate matches, so $metadata.attempts was 1. With the key absent the same call answered at 414 ms. Any session bound below the slowest legitimate answer therefore converts a working call into an unretryable failure, and one above it bounds nothing requestTimeout does not already bound at the stream.

disableConcurrentStreams: true is restated because naming a requestHandler REPLACES bedrock-runtime’s own default provider, which resolves to exactly {disableConcurrentStreams: true} (the defaults mode is legacy, so the mode’s config is empty). Without it the client silently changes connection model, from one isolated session per request to a pooled multiplexed one.

readonly disableConcurrentStreams: true = true;
readonly requestTimeout: 300000 = 300_000;

const STRUCTURED_TOOL_NAME: "emit" = "emit";

The forced-tool name for structured output. One name across every phase, so a decoder can assert on it rather than on positional order in content.

function asResponseBody(payload): InvokeResponseBody;

The parsed payload, read defensively, since every field on the wire is optional.

unknown

InvokeResponseBody


function buildEmbedBody(texts, inputType): string;

The InvokeModel body for one embed request.

output_dimension is named rather than defaulted. Probed live 2026-08-02: the model returns 1536 floats when the field is absent and exactly 1024 when it is present, while the embeddings table stores a fixed-width F32 blob. A default that changed under us would produce vectors of the wrong width against a schema that cannot hold them, and the failure would surface as a distance function returning nonsense rather than an error.

embedding_types: ["float"] is what puts the vectors under embeddings.float; without it the response nests them elsewhere and the reader below finds nothing.

readonly string[]

"search_document" | "search_query"

string


function buildInvokeBody(
key,
prompt,
options,
tool?
): string;

Build the request body in the model’s own dialect. The tool argument selects the lane. When it is absent the model answers in prose. When it is present the request constrains the model to exactly one schema-shaped answer: a forced emit tool call on the Anthropic dialect, a strict json_schema response format on the OpenAI one.

system is omitted rather than sent empty, because an empty system block is a distinct (and rejected) input from no system block at all. An omitted system also has nothing to cache, so cacheSystem over an absent or empty system emits no system key at all instead of an empty cached block.

"sonnet-5" | "opus-5" | "fable-5" | "gpt-5.6-sol" | "gpt-5.6-terra"

string

GenerateOptions

StructuredTool

string


function chunkTexts(texts, size?): readonly readonly string[][];

Slice texts into request-sized chunks, order-preserving. Exported so a test can pin the boundary arithmetic without a client. An off-by-one here drops or duplicates a vector, which lands in the index as a chunk pointing at the wrong body.

readonly string[]

number = EMBED_BATCH_LIMIT

readonly readonly string[][]


function clampTokens(requested): number;

Bound a requested budget to what Bedrock accepts. Above the ceiling the service raises a ValidationException rather than clamping, so an unbounded caller value would fail the call instead of shortening the answer. The floor is 1 for the same reason one axis down: max_tokens must be a positive integer, so a zero or negative request would also fail the call at the service — the floor turns it into the smallest budget the wire accepts, where the truncation gate then reports the response as incomplete rather than the request as malformed.

“Positive integer” is the whole contract, so the clamp also has to answer the two values a min/max pair passes through unchanged. NaN loses every comparison, so it survives both bounds and JSON.stringify writes it as max_tokens: null — the malformed request this function exists to prevent — and it is not a budget at all, so it resolves to the default rather than to a bound. A fractional request IS a budget, so it truncates toward zero and then meets the floor; Infinity truncates to itself and meets the ceiling.

number | undefined

number


function decodeToolInput<A, I>(schema, input): Effect<A, LlmContractViolation>;

Decode a forced-tool payload against its schema.

onExcessProperty: "error" is the option this decode depends on. The default, "ignore", strips an undeclared key and SUCCEEDS (verified against effect 4.0.0-beta.102), which would let a model answer a schema next to the one it was given and have the extra field vanish.

One failure shape is repaired before the violation is constructed: a top-level container field double-encoded as a JSON string (unwrapDoubleEncoded). The repaired payload goes through the SAME strict decode, and a repair that still does not satisfy the schema reports the original payload’s violation, so the repair cannot mask a genuinely off-schema answer.

undefined input means the model produced no emit call at all. That is the same class of failure as a malformed one, and the reason text names it so a caller can tell the two apart in a log without a second error type.

A

I

Codec<A, I>

unknown

Effect<A, LlmContractViolation>


function fromProxyResponse(route, payload): unknown;

Fold a proxy response into the InvokeModel payload the lane expects.

Messages and chat-completions responses are already the payloads asResponseBody and normalizeOpenAiResponse read, so they pass through untouched. An OpenAI embeddings response (data: [{index, embedding}]) becomes Cohere’s embeddings.float, ordered by index because the OpenAI shape permits any order and readEmbeddings pairs vectors with texts positionally. An off-shape embeddings payload folds to an empty float, which readEmbeddings then reports as a count mismatch — the same typed failure a short Cohere answer produces.

ProxyRoute

unknown

unknown


function incompleteReason(parsed): string | null;

The incomplete stop_reason, or null when the response ran to a natural end.

InvokeResponseBody

string | null


function invokeClientFor(config): InvokeClient;

The one construction site for a lane’s transport: the proxy when the configuration names one, Bedrock directly otherwise. Both ModelClientLive and EmbeddingsLive call this, so the two lanes cannot disagree about where a night’s traffic goes.

LlmConfigShape

InvokeClient


function invokeJson(
client,
modelId,
body
): Effect<unknown, ModelUnavailable>;

One InvokeModel round trip: send the body, decode the JSON payload. Both the transport rejection and an unparseable payload land on ModelUnavailable, because neither one says anything about the model’s answer. They say only that no answer arrived. Reading the answer, and judging whether it honors its contract, is the caller’s job.

InvokeClient

string

string

Effect<unknown, ModelUnavailable>


function isRetryableProxyFailure(cause): boolean;

Which failures are worth a second attempt: a throttle, a timeout at the proxy, an upstream that fell over, and a socket that never answered. A 4xx other than those is the request’s fault and repeats identically; an abort is the caller’s decision.

unknown

boolean


function makeBedrockClient(region): BedrockRuntimeClient;

One construction for both lanes: maxAttempts: 10 with adaptive retry absorbs throttles below Effect, and REQUEST_HANDLER_OPTIONS bounds a hung socket.

string

BedrockRuntimeClient


function makeEmbeddings(client): EmbeddingsShape;

The service over an already-built client. Exported as the seam every test uses: a fake InvokeClient records each request body, so the batch boundaries and the wire fields are asserted against the bytes that would go to Bedrock rather than against a mock’s recollection of them.

InvokeClient

EmbeddingsShape


function makeModelClient(client, defaults?): ModelClientShape;

InvokeClient

ModelClientDefaults = {}

ModelClientShape


function makeProxyClient(config, options?): InvokeClient;

The client. send mirrors BedrockRuntimeClient.send closely enough that invokeJson cannot tell them apart: it resolves with the response bytes and rejects with an Error whose name: message modelFailure renders.

Each attempt is bounded by the same per-request inactivity window the Bedrock client uses (REQUEST_HANDLER_OPTIONS.requestTimeout), composed with the caller’s own signal, and a rejected attempt is retried under PROXY_BACKOFF when isRetryableProxyFailure says so. The proxy’s response bytes are returned verbatim for the two chat lanes and re-encoded for the embedding lane after fromProxyResponse.

ProxyConfig

ProxyClientOptions = {}

InvokeClient


function modelByKey(key): ModelInfo;

Resolve a key to its model. Total over ModelKey, so the type makes the throw unreachable. It is there to fail loudly if the table is ever edited out of agreement with the literal union, not to be caught.

"sonnet-5" | "opus-5" | "fable-5" | "gpt-5.6-sol" | "gpt-5.6-terra"

ModelInfo


function modelFailure(modelId, cause): ModelUnavailable;

A Bedrock rejection reduced to the model and the driver’s own summary. The reason carries no prompt and no memory body, because a ModelUnavailable goes back to an agent through a tool response, and the corpus content that produced it does not need to be repeated there.

string

unknown

ModelUnavailable


function normalizeProxyBaseUrl(raw): string;

The origin memhtml will append /v1/... to, or a thrown Error naming what is wrong with it.

Trailing slashes are dropped so http://host:4000/ and http://host:4000 are one value, and the scheme is required because a bare host:4000 is what a shell profile most easily gets wrong and fetch would reject it with a message that does not name the variable. Throwing rather than returning null is deliberate: a set-but-unusable value dying at construction, with the variable named, is the outcome MEMHTML_CONSOLIDATOR_TURN_TIMEOUT_MS chose for the same reason — a typo that silently fell back to Bedrock direct would route production traffic somewhere the operator did not point it.

string

string


function openAiPromptCacheFrom(raw): OpenAiPromptCache;

Parse a raw OPENAI_PROMPT_CACHE_VAR value. Unset or blank is the default, for the reason every proxy variable treats blank so — a blank export is how a variable goes missing. Compared case-insensitively after trimming. Anything else throws with the variable named, matching normalizeProxyBaseUrl: a typo that silently fell back to the default would put the write premium back on every call while the environment claimed to have turned it off.

string | undefined

OpenAiPromptCache


function parseProxyModelMap(raw): ReadonlyMap<string, string>;

Parse PROXY_MODEL_MAP_VAR. Empty entries (a trailing comma) are ignored; an entry with no =, an empty side, or a key mapped twice throws with the offending entry quoted, for the reason normalizeProxyBaseUrl records.

string

ReadonlyMap<string, string>


function proxyConfigFromEnv(env?): ProxyConfig | null;

The proxy configuration an environment describes, or null when PROXY_BASE_URL_VAR is absent or blank — which is the default, Bedrock direct.

A blank value reads as absent for the reason the credential preflight treats "" as unset: a blank export is how a variable goes missing in practice, and an empty origin is not a place to send traffic.

Record<string, string | undefined> = process.env

ProxyConfig | null


function proxyModelId(config, modelId): string;

The id the proxy is asked for: the map’s exact name when it has one, else the prefix and the Bedrock id. The map wins so one odd model can be named by hand while the rest follow the rule.

ProxyConfig

string

string


function proxyModelPrefix(raw): string;

The prefix a raw PROXY_MODEL_PREFIX_VAR value means: unset or blank is the LiteLLM default, PROXY_MODEL_PREFIX_NONE is no prefix at all, anything else is taken as written after trimming. Blank reads as unset for the reason every other variable here treats it so — a blank export is how a variable goes missing — and because Config cannot see the difference.

string | undefined

string


function proxyRouteFor(modelId): ProxyRoute | null;

The route for a model id, or null for an id this package does not call. Anthropic models speak Messages, the OpenAI model speaks chat completions, and the one embedding model is its own lane.

string

ProxyRoute | null


function readEmbeddings(payload, expected): ModelUnavailable | readonly Float32Array<ArrayBufferLike>[];

Read the vectors out of a decoded payload, or say why they are unusable.

A count mismatch is a typed failure rather than a short array, because the caller pairs vectors with chunk ids positionally. A response one vector short would shift every subsequent pairing and store each embedding against the wrong body. A width mismatch fails the same way one axis over. The embed_model watermark records id@dim, so a vector of another width cannot be compared against the stored ones.

unknown

number

ModelUnavailable | readonly Float32Array<ArrayBufferLike>[]


function readText(parsed): string;

Join the text blocks. Thinking blocks are discarded on purpose, because a caller reads the answer rather than the deliberation. Concatenating the two would put reasoning the model did not commit to into the value a phase acts on.

InvokeResponseBody

string


function readToolInput(parsed): unknown;

The forced tool’s input, or undefined when the model answered without calling it. Matched on the block’s name instead of its position, because a thinking block precedes the tool call on the two adaptive models and an index-based read would find that block.

InvokeResponseBody

unknown


function thinkingFor(key):
| {
type: "adaptive";
}
| null;

The thinking object per Anthropic model. Opus 5 and Fable 5 take {type: "adaptive"} (Fable is adaptive-only). Sonnet 5 reasons unconditionally and takes NO thinking key. Sending one to Sonnet 5 raises a validation error instead of being ignored. The OpenAI lane never consults this: its reasoning dial is reasoning_effort alone.

Verified live 2026-08-02: all three Claude models accept this shape alongside a forced tool_choice, so structured output and adaptive thinking compose.

"sonnet-5" | "opus-5" | "fable-5" | "gpt-5.6-sol" | "gpt-5.6-terra"

{
type: "adaptive";
}
readonly type: "adaptive";

The only thinking mode the Claude 5 models accept here.


null


function toInputSchema(schema): JsonSchemaObject;

Derive the tool’s input_schema from an effect schema.

Schema.toJsonSchemaDocument returns its definitions map separately from the schema while emitting $ref: "#/$defs/<name>" into the schema, so the map is folded back under the root as $defs — the pointer those refs already name. Verified live 2026-08-02 that Bedrock resolves a $ref into a root-level $defs inside input_schema.

A RECURSIVE schema is what makes this load-bearing, and it is the only thing that does: probed 2026-08-25 on effect 4.0.0-rc.111, a struct reused from two places is emitted inline at each, while a self-reference is hoisted because it has no finite expansion. So the fold is a narrow repair for recursion rather than the general case, and definitions is usually empty.

A numeric field should be declared Schema.Finite, not Schema.Number: the latter emits an anyOf with a string branch for Infinity/NaN, which invites the model to answer a number field with the string "NaN".

Top

JsonSchemaObject


function toProxyRequest(
config,
modelId,
body
): ProxyRequest | null;

Translate one InvokeModel request into the proxy’s request for it. Pure, so the wire test pins every field against the bytes that would go out.

model is the proxy’s name for the Bedrock id (proxyModelId: bedrock/<id> by default, or the map’s exact name).

  • Messages: anthropic_version is Bedrock’s field and the Messages API has no such key, so it is dropped; model is added. Everything else — max_tokens, system (with any cache marker), thinking, output_config, tools, tool_choice — is already the Messages dialect.
  • Completions: the body is already chat completions; only model is added.
  • Embeddings: Cohere’s texts becomes input, output_dimension becomes dimensions, and input_type rides through under its own name — the proxy forwards it to Cohere, and it is the field the retrieval asymmetry depends on (embeddings.ts). encoding_format: "float" asks for the plain arrays the fold below reads.

ProxyConfig

string

Record<string, unknown>

ProxyRequest | null


function wrapAsData(label, text): string;

Wrap rollout or memory text for a user turn. The delimiters keep the content’s own prose from being read as a directive to the model. That prose is often instruction-shaped in this corpus, because the memories record instructions.

The delimiters only hold if the content cannot produce them: a body carrying the literal closing tag would end the data block early and place its own remainder OUTSIDE the boundary, where it reads as the caller’s instructions. So every end tag closingTagsFor names is neutralized in the content before wrapping — the slash gains a backslash, which keeps the text legible, and the attribute text if any is kept, so the neutralizer rewrites nothing but the one character that made the text a delimiter. Case-insensitive, because the boundary is prose to the model rather than parsed markup, and a cased variant reads as the same tag.

string

string

string