How Workstate works
Workstate does three jobs. It indexes your sources, it answers searches with citations, and it records what agents write to the ledger. This page follows your data through each job.
The three flows
text
Indexing, on every sync
source --> read --> filter --> compare --> chunk --> embed --> search index
Searching, on every query
question --> embed --> nearest chunks --> rerank --> results with citations
(one namespace only)
Recording, on every ledger write
agent --> check key --> ledger (readable at once) --> search index (within seconds)Everything happens inside one namespace: a separate index with its own sources, search results and ledger.
Indexing: from a source to the index
A sync runs on a schedule, when someone selects Sync now, when an agent calls reindex_corpus, and after every upload. Each sync takes these steps:
- Read the source. For GitHub, Workstate reads the latest commit on the default branch of each shared repository. For Confluence, it reads the current pages of the chosen spaces. For Jira, it reads the issues and their comments. For uploads, it reads the files you uploaded, and extracts the text from PDF and Office files.
- Filter the files. It skips hidden files and folders (such as
.envand.git), dependency and build folders (such asnode_modules,vendoranddist), lockfiles, minified files, binary files, and text files over 1 MB. - Compare with the last sync. Only new and changed files go further. Files that are gone from the source are removed from the index. If the source could be read only in part, nothing is removed.
- Split each file into chunks. A chunk is a passage of a file, and it is the unit that search returns. Code is split along its syntax for Rust, Python, TypeScript, JavaScript, C, C++, Bash and Ruby. Markdown is split along its headings. Other text is split into passages of similar size.
- Embed each chunk. An embedding model turns the chunk's text into a vector: a list of numbers that represents its meaning. Passages with similar meanings get similar vectors, so search can match meaning rather than exact words.
- Store the chunks in the search index. Each chunk is labelled with its namespace, corpus, repository, file path and line range. That label becomes the citation.
Each file lands in one corpus. All GitHub files go to code, including Markdown. Confluence pages, Jira issues and uploaded files that are not code go to wiki. Uploaded code goes to code. See Sources & sync.
Searching: from a question to cited results
When an agent calls search_corpus, or a person uses Search in the console, Workstate takes these steps:
- Settle the namespace. Workstate uses the namespace the request names, if the person making the request can reach it, and refuses the request otherwise. A request that names no namespace uses the person's oldest one. A search never reaches beyond one namespace.
- Embed the question with the same model, so the question and the chunks can be compared.
- Find the nearest chunks in the namespace, applying any filters: corpus, repository, language or path.
- Rerank them. A second model reads the question together with each candidate, and scores how well the candidate answers it. This step is slower than the first pass, and more accurate.
- Return the best results, 8 by default. Each result carries its provenance: repository, file path, line range and scores. Ledger results also carry the topic id, type and status, and each topic appears only once.
See Search & citations.
Recording: from an agent to the ledger
When an agent calls ledger_create or ledger_append, Workstate takes these steps:
- Check the API key. The gateway works out the person from the key, and checks the requested namespace against that person's access. The key's owner becomes the author, and the agent label comes from the call.
- Store the entry in the ledger.
ledger_getand the topic's page in the console show it at once. - Index the entry for search within a few seconds. Workstate indexes the topic's title, summary and current state, and each history entry on its own, so older reasoning stays findable.
The ledger is the one part of Workstate that is not a copy of something you keep elsewhere. It holds what your team learned while working. See The ledger.
How agents and people connect
Agents connect to two MCP (Model Context Protocol) servers over Streamable HTTP, the MCP transport for remote servers:
| Server | Address | Tools |
|---|---|---|
rag-corpus | https://mcp-uat.workstate.io/corpus | search_corpus, reindex_corpus |
rag-ledger | https://mcp-uat.workstate.io/ledger | ledger_search, ledger_get, ledger_create, ledger_append, ledger_move, ledger_archive |
Every request carries the person's API key in the Authorization header. A request can also name a namespace in the X-RAG-Namespace header, and label the agent in the X-RAG-Agent header. See Connect your agent and MCP tools.
People use the console. They sign in with GitHub, search, browse the ledger, and manage sources, API keys, the team and namespaces. The console shows the ledger read-only, because agents write to it.
Your own software can call the HTTP API that the console uses, with an API key. A public API reference is not published yet Coming soon. See Your own agent.
Where the models run
Today, Workstate embeds and reranks your content with models that the Workstate deployment runs itself. If that changes, Security & privacy will list any sub-processor before one is used.