Skip to content

@memhtml/traces

@memhtml/traces: a streaming JSONL parser and trace indexer over agent session transcripts. Read-only, and it stores no session content.

The package turns a directory of session transcripts into rows the index can hold: pointers, counters, and capped text heads, never the transcript itself. discover names the layout of that directory (a projects/ tree of per-cwd slug directories, each session a <sessionId>.jsonl, each subagent an agent-<agentId>.jsonl sidecar under subagents/). watermark decides, from a file’s size, modification time and a stored byte offset, whether a scan can skip the file, read only its tail, or must rescan it from byte 0. parse streams one file line by line, splitting on the raw buffer so an unterminated trailing line is neither folded nor counted as consumed. extract folds those lines into a SessionExtract without ever throwing, because another process may be appending to the file mid-line. scan composes the four into the unit the indexer persists and owns mergeTailExtract, the rule for folding a tail’s extract into a stored row.

Persistence is not here. @memhtml/index’s trace persister owns the traces and trace_watermarks tables and binds to these shapes structurally; the CLI’s trace operations call scanTraceRoot and mergeTailExtract from this package and hand the results to that persister.

The package publishes one import path, @memhtml/traces, which re-exports the five modules below, so the reference on this page is the whole exported surface.

Module What it holds
discover.ts The transcript-tree layout: PROJECTS_DIR, SUBAGENTS_DIR, and the walk that yields each session file with its sidecars.
watermark.ts Watermark, WatermarkAction, watermarkPlan and advanceWatermark: the pure skip, tail or rescan arithmetic over a file’s stat.
parse.ts The streaming reader that splits a transcript into lines and reports what one scan consumed.
extract.ts READ_RECORD_TYPES, the per-line fold (emptyAccumulator, foldLine, finalizeExtract) and the SessionExtract it produces.
scan.ts scanTraceRoot, WatermarkReader, the per-file outcome, and mergeTailExtract.

Mutable fold state. Private, so callers see SessionExtract only.

readonly agentIds: Set<string>;
aiTitle: string | null;
cwd: string | null;
droppedLines: number;
droppedNoSession: number;
entrypoint: string | null;
firstPrompt: string;
gitBranch: string | null;
maxEpochMs: number | null;
maxIso: string | null;
minEpochMs: number | null;
minIso: string | null;
readonly modelCounts: Map<string, number>;
parsedLines: number;
readonly prompts: Map<string, {
hasText: boolean;
row: PromptRow;
}>;
sessionId: string | null;
skippedTypeLines: number;
turnCount: number;
unknownTypeLines: number;
version: string | null;

How the reader learns a file’s slug when the caller has not already discovered it.

readonly slug: string;

A file’s current stat, the half of Watermark the filesystem supplies.

readonly mtimeMs: number;
readonly size: number;

How many lines each disposition claimed. Every line lands in exactly one counter, so parsedLines + droppedLines is the number of non-empty lines the scan consumed.

  • parsedLines: decoded to a JSON object, whatever happened afterwards.
  • droppedLines: undecodable JSON, or decoded to something that is not an object with a string type. Counted rather than fatal, so one bad line does not abandon the file.
  • droppedNoSession: a read-type record with no string sessionId. Subset of parsedLines.
  • skippedTypeLines: a SKIP_RECORD_TYPES type. Subset of parsedLines.
  • unknownTypeLines: a type in neither list, meaning a record type the runtime added after this allowlist was written. Subset of parsedLines, and the counter to watch when a new Claude Code release lands.
readonly droppedLines: number;
readonly droppedNoSession: number;
readonly parsedLines: number;
readonly skippedTypeLines: number;
readonly unknownTypeLines: number;

What one file’s scan consumed, alongside the extract.

readonly bytesRead: number;

Bytes occupied by newline-terminated lines, terminators included. The next read starts at startByte + bytesRead, so this is what keeps a tail landing on a record boundary.

readonly extract: SessionExtract;
readonly readFailed: boolean;

True when the stream itself could not be read: an absent file, a permission rejection, a transient IO error. The extract is empty and bytesRead is 0 however far the read got before breaking, because a partial fold ends on no boundary a watermark could stand on.

A caller holding a watermark must leave it untouched on a failed read. Stamping the file’s current size and mtime over a read that consumed nothing makes the next run’s comparison see an unchanged file and skip it, which turns one transient error into a transcript that is never indexed.

readonly startByte: number;

Byte offset the read began at, 0 for a full scan.


One trace_prompts row (design §3.3), minus the session_id the enclosing SessionExtract carries.

readonly agentId: string | null;

Set only on a subagent sidecar’s records.

readonly at: string;

ISO-8601 UTC instant of the prompt’s first record.

readonly ordinal: number;

0-based position of this prompt among the distinct prompts of this session, in first-appearance order. Per-session scope, so it is comparable only within one session_id.

readonly promptId: string;
readonly textHead: string;

First TEXT_HEAD_LIMIT characters of the prompt’s text, whitespace-collapsed.

readonly turnUuid: string;

The uuid of the first user record carrying this promptId, the (sessionId, uuid) cite.


One file’s outcome, whether or not it was read.

readonly action: WatermarkAction;
readonly agentCount: number;

agent_count for the traces row: this file’s distinct agentIds unioned with the sidecar filenames of its session. 0 for a skipped or unreadable file.

readonly extract: SessionExtract | null;

Absent for a skip, because a skipped file is not opened and yields no extract. Also absent when the read failed, so a persister writes nothing for the file and the next run retries it.

readonly file: SessionFile;
readonly watermark: Watermark;

The watermark to store. Unchanged from the previous one for a skip and for a failed read, because a failed read consumed nothing and advancing past the stored offset would make the next run’s size+mtime comparison skip a transcript nobody has read.


A whole scan: per-file outcomes plus the totals an operator reads.

The four action counters PARTITION files: skipped + tailed + rescanned + failed is exactly files.length. That is why a failed read counts here and nowhere else — counting it as tailed would report a transcript as read when its rows were never extracted, and a scan whose numbers add up is the only way an operator can tell a quiet night from a broken one.

readonly bytesRead: number;

Bytes actually read. The number the incremental design exists to keep small.

readonly failed: number;

Files the scan planned to read and could not: an absent file, a permission rejection, a transient IO error. Each holds its stored watermark, so the next run retries it.

readonly files: readonly ScannedFile[];
readonly rescanned: number;

Files read from byte zero, and read successfully.

readonly skipped: number;
readonly tailed: number;

Files read from their recorded offset forward, and read successfully.


Everything one session file yields: the traces row’s content fields, its prompt rows, and the scan’s counters. file_size/file_mtime/indexed_at are absent. They belong to the stat the watermark already took, and duplicating them would give the same fact two sources.

readonly agentIds: readonly string[];

Distinct agentId seen in this file, first-appearance order. Empty for a main session.

readonly aiTitle: string | null;
readonly counters: ParseCounters;
readonly cwd: string | null;
readonly endedAt: string | null;

Latest record instant, ISO-8601 UTC.

readonly entrypoint: string | null;
readonly filePath: string;

Absolute path of the file scanned.

readonly firstPrompt: string;
readonly gitBranch: string | null;
readonly model: string | null;

Most frequent message.model on assistant records, excluding SYNTHETIC_MODEL.

readonly promptCount: number;

Distinct promptId count on user records. Equals prompts.length.

readonly prompts: readonly PromptRow[];
readonly sessionId: string | null;

From the first record carrying one, bare types included. null when the file has none.

readonly slug: string;

The ~/.claude/projects/<slug> directory name, which is a path slug. The slug field on a subagent record is a title slug for the agent’s task, and is a different fact entirely.

readonly startedAt: string | null;

Earliest record instant, ISO-8601 UTC.

readonly turnCount: number;

Enveloped records, meaning those carrying a uuid. A pr-link has a session but no turn.

readonly version: string | null;

One discovered transcript file with the stat the watermark decision needs.

readonly agentId: string | null;

Set only on a sidecar, from its agent-<agentId>.jsonl filename.

readonly filePath: string;

Absolute path, built from the caller’s traceRoot.

readonly kind: SessionFileKind;
readonly mtimeMs: number;
readonly sessionId: string;

The session this file belongs to, read from the path. That is the filename stem for a main session and the owning directory name for a sidecar. Deriving it from the path costs one readdir per directory and opens no file.

readonly size: number;
readonly slug: string;

The projects/<slug> directory name, which is the cwd slug, traces.slug.


What a previous scan recorded about a file, mirroring the trace_watermarks row (design §3.3). @memhtml/index’s trace persister owns the table; this module owns the arithmetic.

  • size: file length in bytes at scan time.
  • mtimeMs: modification time in milliseconds since the Unix epoch, the unit node:fs’s Stats.mtimeMs reports. The SQL column stores an ISO-8601 string, so the adapter converts at the boundary and this type carries one unit only.
  • byteOff: a 0-based byte offset into the file, one past the last byte consumed, and therefore the start of the next read. Equal to size after a complete scan.
readonly byteOff: number;
readonly mtimeMs: number;
readonly size: number;

An action plus the byte offset to open the read stream at.

readonly action: WatermarkAction;
readonly startByte: number;

0-based byte offset to pass as createReadStream({ start }). Always 0 for a rescan.

type ReadRecordType = typeof READ_RECORD_TYPES[number];

type SessionFileKind = "session" | "subagent";
  • session: <root>/projects/<slug>/<sessionId>.jsonl, the main transcript.
  • subagent: <root>/projects/<slug>/<sessionId>/subagents/agent-<agentId>.jsonl. Its records carry the parent sessionId too, so it indexes into the same session row.

type SkipRecordType = typeof SKIP_RECORD_TYPES[number];

type WatermarkAction = "skip" | "tail" | "rescan";
  • skip: nothing changed, so do not open the file.
  • tail: the file grew by append, so read from WatermarkPlan.startByte.
  • rescan: the file was rewritten, compacted, or is unknown, so read from byte 0.

type WatermarkReader = (filePath) => Effect.Effect<Watermark | null, StorageFailure>;

Reads a file’s stored watermark. null for a file never scanned.

string

Effect.Effect<Watermark | null, StorageFailure>

const FIRST_PROMPT_LIMIT: 500 = 500;

traces.first_prompt is an index entry, not a copy of the prompt.


const PROJECTS_DIR: "projects" = "projects";

The projects/ directory the per-cwd slug directories live under.


const READ_RECORD_TYPES: readonly ["user", "assistant", "system", "attachment", "agent-name", "ai-title", "pr-link"];

Record types the parser reads. Applied as an allowlist before any field access, because the bare types carry no envelope. Reaching for record.cwd on a file-history-snapshot would read a field that does not exist, on a record that is not about a session at all.


const SKIP_RECORD_TYPES: readonly ["last-prompt", "mode", "permission-mode", "queue-operation", "file-history-snapshot", "file-history-delta"];

Counted and skipped, never an error. file-history-snapshot and file-history-delta carry no sessionId and no envelope at all (probed on 5,387 real files, 2026-08-01: their keys are {isSnapshotUpdate, messageId, snapshot, type} and {backup, messageId, snapshotMessageId, trackingPath, timestamp, type}), so a record with no session key is dropped rather than treated as malformed input.


const SUBAGENTS_DIR: "subagents" = "subagents";

The subdirectory holding a session’s subagent sidecars.


const SYNTHETIC_MODEL: "<synthetic>" = "<synthetic>";

The placeholder model id the runtime emits for a non-model turn; never a session’s model.


const TEXT_HEAD_LIMIT: 200 = 200;

trace_prompts.text_head is an index entry, not a copy of the prompt.

function advanceWatermark(
stat,
startByte,
bytesRead
): Watermark;

The watermark to store after a scan consumed bytesRead bytes starting at startByte. size/mtimeMs come from the stat taken before the read, so a file appended to during the scan compares unequal next time and gets tailed rather than skipped.

FileStat

number

number

Watermark


function agentCountFor(extract, sidecarAgentIds): number;

The agent_count for a traces row: distinct agents named by the session’s records unioned with those named by its sidecar filenames (design §7). The union covers both sides, because an agent’s sidecar exists before its first record lands, and a resumed session’s records can name an agent whose sidecar has been pruned.

SessionExtract

readonly string[]

number


function discoverSessions(traceRoot): Effect<readonly SessionFile[], StorageFailure>;

Every transcript file under traceRoot, main sessions and subagent sidecars alike.

traceRoot is a parameter with ~/.claude as the caller’s default, and it is not hardcoded here, so a test drives a fixture tree and an operator can point the indexer at an archive.

A file that vanishes between readdir and stat is dropped rather than failed. Transcripts are written by a live process that may compact or delete one mid-scan, and the next run rediscovers whatever is there.

string

Effect<readonly SessionFile[], StorageFailure>


function emptyAccumulator(): Accumulator;

A fold state with nothing seen yet.

Accumulator


function extractFromText(text, file): SessionExtract;

Fold a whole JSONL text. The streaming reader in parse.ts is the production path; this is the same fold for a string in hand.

string

string

string

SessionExtract


function finalizeExtract(accumulator, file): SessionExtract;

Close the fold into the immutable extract. Pure with respect to the accumulator.

Accumulator

string

string

SessionExtract


function foldLine(accumulator, line): Accumulator;

Fold one raw line into the accumulator. Total, so every input either updates state or a counter, and no input throws. A blank line is not a record and is not counted.

Mutates and returns accumulator. A fresh object per line would allocate once per line of a 3.67 GB corpus.

Accumulator

string

Accumulator


function mergePrompts(stored, tail): readonly PromptRow[];

Concatenate two prompt lists into one per-session first-appearance ordering.

A tail’s ordinals are 0-based over the appended slice, so they are renumbered from the end of the stored list. Without that, every tail would collide with ordinal 0 and trace_prompts.ordinal would stop being an order at all. A promptId present in both sides keeps its stored ordinal, uuid, and instant, since it began before the tail. It takes the tail’s textHead only when the stored one is empty, which is how a prompt whose text arrived after the boundary gets indexed.

readonly PromptRow[]

readonly PromptRow[]

readonly PromptRow[]


function mergeTailExtract(stored, tail): SessionExtract;

Merge a tail’s extract into the session’s stored one. Tails only, because a rescan’s extract already describes the whole file and replaces the stored row outright.

The merge exists because a tail’s extract describes the appended slice and not the session. Its first_prompt is a prompt from the middle of the conversation, its started_at is an hour after the session began, its turn_count counts only new turns, and its prompt ordinals restart at 0. Every field below states which side owns it and why. The producer owns these reading semantics, so an indexer that merged the fields itself would have to rediscover all of it.

SessionExtract

SessionExtract

SessionExtract


function parseSessionFile(
filePath,
startByte?,
identity?
): Effect<ParseResult, never>;

Stream a transcript from startByte and fold it into a SessionExtract.

Cannot fail. A truncated line and a line of binary garbage degrade to counters on the extract, while a read the filesystem refuses outright (absent file, permission rejection, transient IO error) degrades to an empty result flagged readFailed. This runs over thousands of files written by a live process, so one unreadable transcript costs that transcript’s rows and not the whole run. The counters and the flag are what an operator reads in place of an error, and the flag is what lets a caller hold its watermark still so the next run retries.

startByte is a 0-based byte offset and must be one watermarkPlan produced. A caller-invented offset can land mid-line, and that first partial line is then counted as malformed rather than recovered.

string

number = 0

FileIdentity

Effect<ParseResult, never>


function scanTraceRoot(traceRoot, readWatermark): Effect<ScanReport, StorageFailure>;

Scan every transcript under traceRoot, reading only what the watermarks say changed.

traceRoot is a parameter. ~/.claude is the caller’s default rather than this module’s constant, so the whole scan is drivable against a fixture tree.

Files are processed sequentially. Concurrency here would trade a bounded, predictable IO profile for contention with the live process that is writing these transcripts, and the incremental watermark has already reduced a daily run to the handful of files that changed.

string

WatermarkReader

Effect<ScanReport, StorageFailure>


function sessionIdFromPath(filePath, traceRoot): string | null;

The session id a transcript path names, or null when the path is not one. Pure.

string

string

string | null


function sidecarAgentIds(files, sessionId): readonly string[];

The agentIds of a session’s sidecars, from filenames alone. Feeds agent_count without opening a sidecar, because the count is a property of the tree. The sidecars’ own records are indexed on their own pass.

readonly SessionFile[]

string

readonly string[]


function slugFromPath(filePath): string;

The projects/<slug> directory a transcript sits under, derived from its path.

A sidecar sits two levels deeper (<slug>/<sessionId>/subagents/), so the slug is the grandparent’s parent there. Returns "" for a path outside the tree. The extract still carries a real session_id from the records, so an unslugged file indexes rather than being lost.

string

string


function userText(message): string;

The text of a user record’s message.content, or "" when it carries none.

Content arrives in two shapes and both hold real prompt text: a bare string, and a block list whose text blocks are joined. Probed 2026-08-02: of the distinct prompts in six large sessions, the block-list form was the first appearance for 24 of 30 in one file and 34 of 38 in another, so a string-only rule would leave first_prompt empty for most sessions.

A tool_result-only list yields "", because a tool’s output is not something the user said.

unknown

string


function watermarkAction(prev, curr): WatermarkAction;

Decide how to read a file given what the last scan recorded.

Both size and mtime must match to skip. Size alone would miss an in-place rewrite that happens to preserve the length; mtime alone would miss a write inside the same clock tick.

A grown file is tailed only when mtime also advanced or held steady. A file that grew while its mtime moved backward was restored or rewritten instead of appended to, so it is rescanned. Shrinking is unambiguous, because bytes the watermark counted are gone and any offset into the file is now meaningless.

A byteOff past the current size is treated as a rescan even when size grew, because that offset could only come from a larger earlier file. The growth is a rewrite that has not yet reached the old length.

Watermark | null

FileStat

WatermarkAction


function watermarkPlan(prev, curr): WatermarkPlan;

watermarkAction with the read offset the caller needs.

Watermark | null

FileStat

WatermarkPlan