---
title: memhtml
description: "An agent's long-term memory as a git repository of semantic HTML files, with a rebuildable SQLite index, retrieval that fuses four ranking arms, and a seventeen-phase curation pass."
---

<div class="rfc-title-block">
  <p class="rfc-memo-title">Memory for Agents, in HTML</p>

  <p class="rfc-brand">MEANING · MEMORY · MARKUP</p>
</div>

## What memhtml is

memhtml stores an agent's long-term memory as a git repository of small HTML files, one fact per
file. A SQLite index over that repository makes the files searchable, and an agent reads and writes
the store either through the `memhtml` command line or through an MCP server that serves the same
store over stdio. What the agent remembers is a directory on disk that you can open in a browser,
diff, and edit with ordinary tools.

## The problem it solves

An agent that writes down what it learns needs three answers later: what do I know, where did it
come from, and what has changed since. Keeping those facts as rows in a vector database answers the
first one well and drops the other two, because a row carries no history and a re-embedding pass
rewrites the store with nothing for a reviewer to read.

memhtml answers all three out of git. Each fact is a file, so you read it without a query. Each
change to the corpus is exactly one commit, so `git log` in the store is the history of what the
agent learned, and a night of automated curation is a diff a person can review. Retiring a memory
moves the file into `archive/<YYYY>/` with the original path mirrored beneath it, so
`git log --follow` reads through a memory's whole life and no correction destroys the fact it
corrected. The index holds no facts of its own: delete `.memhtml/index.db` and
`memhtml index rebuild` reconstructs it from the tree.

## What it is made of

Each memory is standard HTML5 that a browser renders with no build step. The `<meta>` elements in
the head carry the typed metadata, one `<mark>` span in the body carries the claim the file exists
to state, and `<link>` elements carry the relationships an author asserted between memories. One
fact per file is what keeps a correction reviewable: `memhtml correct` writes the replacement and
archives the original in one commit, so an interrupted run leaves no pair of live contradicting
memories behind.

Two SQLite databases sit under `.memhtml/`, opened through Node's built-in `node:sqlite` driver, and
they differ in what a rebuild can recover. `index.db` holds the projected rows, the full-text index,
the embedding vectors, and the mined edges, all of which the indexer regenerates from the tree.
`state.db` holds the facts the tree cannot reproduce: access counts, reinforcement counts, and
outcome scores. Its durable copy is a committed file, `.memhtml/state/access.jsonl`.

Retrieval runs four independent ranking arms over one query and then combines their orderings.
Full-text search ranks on title, claim, and body text. Vector similarity ranks on the embeddings.
Recency ranks an episodic memory by when the fact happened rather than by when someone wrote it
down. Salience ranks on how often someone has chosen to open the memory. One SQL statement fuses the
four orderings by reciprocal rank fusion, where each arm contributes `1/(rank + 60)` under its own
weight, and a pass in TypeScript then drops near-duplicate hits. That fusion is what this site
means by four-arm RRF. When an arm's precondition is missing, memhtml drops that arm and sets
`degraded: true` on the response, so the search gets narrower and still answers. A store with no
embedder bound is the common case, and it still returns ranked hits from the other three arms.

Three doors write into the same tree: `memhtml write` and `memhtml apply` on the command line, the
write tools on the MCP server, and your own file tools against the checkout. The store owns git
staging, so one call becomes exactly one commit and a caller cannot bundle two unrelated writes into
it.

`memhtml sleep run` is the curation pass you schedule overnight. It walks seventeen ordered phases on
a `sleep/<date>` branch and commits each phase's work on its own, so a reviewer reads the night one
phase-shaped diff at a time. It detects contradictions and leaves the choice of a winner to a writer
or a human, because that choice is a one-way door. `memhtml sleep merge` then refuses to move `main`
when the night made retrieval worse.

Every command writes one JSON envelope to stdout and nothing else, sends its logs to stderr, and
exits 0 for success, 2 for a usage error, and 1 for a runtime failure. `memhtml manifest` prints
every command, flag, response type, error code, and environment variable the binary accepts, and it
answers on a machine with no store, no database, and no credentials.

## Where to start

Install the binary and work through the four tutorials in order. One published package,
[`memhtml`](https://www.npmjs.com/package/memhtml), carries the whole system, and a clone-and-build
path exists for working on it. If you are an AI agent rather than a person,
read [For agents](/agents/) first: it states the assumptions to drop and the shortest path to a
correct first call.

<div class="rfc-toc">
  * **1. [Learn](/learn/)**: from a clone to a working store, then task-shaped how-tos
    * 1.1. [Install memhtml and initialize a store](/learn/tutorial/install/)
    * 1.2. [Write your first memory](/learn/tutorial/first-memory/)
    * 1.3. [Retrieve it](/learn/tutorial/first-retrieval/)
    * 1.4. [Wire up the MCP server](/learn/tutorial/mcp-server/)
    * 1.5. [Run the store day to day](/learn/operations/run-the-store-day-to-day/)
  * **2. [Reference](/reference/)**: generated from the registries the binary parses its own arguments with
    * 2.1. [Global flags](/reference/global-flags/) and [configuration](/reference/config/)
    * 2.2. [Response types](/reference/response-types/) and [error codes](/reference/error-codes/)
    * 2.3. [MCP tools](/reference/mcp-tools/) and [resources](/reference/mcp-resources/)
    * 2.4. [Vocabularies](/reference/vocabulary/), [sleep phases](/reference/sleep-phases/), [RRF arms](/reference/rrf-arms/)
    * 2.5. [Index schema](/reference/schema/) and the [requirements ledger](/reference/requirements/)
    * 2.6. [Contracts](/reference/contracts/): every exported symbol of `@memhtml/contracts`, with its doc comment
  * **3. [Internals](/internals/)**: why the system is shaped this way, and what each decision refuses
    * 3.1. [Packages and dependency direction](/internals/packages-and-dependency-direction/)
    * 3.2. [The memory file format](/internals/the-memory-file-format/)
    * 3.3. [The write path](/internals/the-write-path/)
    * 3.4. [Four-arm retrieval](/internals/four-arm-retrieval/)
    * 3.5. [The sleep pipeline](/internals/the-sleep-pipeline/)
    * 3.6. [Testing posture](/internals/testing-posture/)
  * **Appendix A. [Glossary](/glossary/)**: the domain vocabulary, each term linked to the chapter that develops it
</div>

## How to read this site

The Reference tier is generated from the same registries the binary parses its own arguments with,
so a flag, error code, or response type on those pages is the one the binary ships. Where this site
and the binary disagree, the binary is right and a test is missing.

Architectural claims cite the code in repo-relative `path:line` form. Those line numbers point into
the commit this site was built from, so read them as a pointer into the source rather than as a
permanent address. Every measurement on the site carries the date someone took it.

Section numbers live in the heading text, as in `## 3.2. Edge encoding`, so the page, the contents,
the search index, and the Markdown this site serves to an agent all agree on which section §3.2 is.
Anchors stay unnumbered, so inserting a section does not break an inbound link.

Every page also serves its own Markdown source at the same path with `.md` appended, and links that
twin from its `<head>`. [`llms.txt`](/llms.txt) indexes the whole corpus for an agent that would
rather read the document than the page.