(CASE STUDY · 02)
OREAG
- RAG
- Agent Memory
- MCP
- pgvector
- FastAPI · Next.js
Oreag is a RAG and memory service. A developer uploads documents, tunes chunking, the embedding model and the answer model with their own provider keys, and gets a per-project API that answers with the file, page and similarity of every source. Coding agents use the same project as memory, through an MCP server.
- ROLE
- Product & Engineering
- TIMELINE
- 7 Months
- YEAR
- 2026
- TEAM
- Solo
(MY ROLE)
- Designed and built Oreag alone: the Next.js dashboard, the FastAPI backend, the pgvector schema and the MCP server.
- Built a read path that searches two ways and asks back when the evidence is thin.
- Built usage metering, a query log and a Gaps inbox that turns flagged questions into evaluator tests.
- Wrote the in-app docs, a C4 model and a CI check that fails when either drifts from the code.
- (WHY IT EXISTS)
- An app needs answers from its own documents. A coding agent needs a place to remember. Neither should mean running another vector database, so one Oreag project serves both.
- (THE FIRST VERSION WAS A SCRIPT)
- February 2026: a LangChain script. Text files in, 1000-character chunks with 200 overlap, OpenAI embeddings, a Chroma collection. The chunk defaults survived. Chroma did not.
- (WHAT CHANGED)
- A deploy hook failed every file in flight, so the files table became a queue with leases. Keyword search matched strings, not words, so it now stems. The semantic cache matched topics, so its floor rose from 0.75 to 0.90.
Process
-
01
Prototype & Research
A 36-line LangChain indexer in February, a LangGraph reference in April. On June 11 the script moved to legacy/ and the platform build began.
-
02
Platform & Keys
The dashboard and API landed on June 12, bring-your-own keys on June 17, encrypted with Fernet. Each project got its own /v1 endpoint and oreag_sk_ key.
-
03
Memory & Agents
Agent memory landed on June 18: a memories table, an MCP server and a Memory tab. Memories shared the chunks' vector space, and the 3D graph followed on July 3.
-
04
Retrieval & Trust
Hybrid search and the semantic cache landed on July 5, stemming on August 13. September added the evaluator, Queries, Health and Gaps, to find where answers fall short.
How It Works
Upload a file and a worker converts, chunks and embeds it into pgvector. Ask a question and the API checks its cache, searches two ways, and answers only when the evidence holds. Agents reach the same project through MCP.
The Workspace
Every project is one page with six tabs: Files, Memory, Playground, API, Visualize and Settings. Files shows what is indexed and which edition is in force. The Playground runs the same code path as the public API, so an answer here is an answer there. Both screens show sample data from Oreag's own docs.
In The Detail
Every answer carries its sources: file, page and match score, with the cited ones marked. Agents keep notes in the same project, and pinned ones come first in the list an agent reads when a session starts.
Under The Hood
Bring your own key for any of 16 hosted providers, or run one of 3 local runtimes with no key. Keys are encrypted at rest and resolved per request: the project's own key first, then the account's. A new project starts from the defaults below.
(SCOPE)
- file extensions on the upload allowlist
- 29
- model providers, keyed or local
- 19
- MCP tools for coding agents
- 9
THE DEFAULTS
- Chunking
- 1000 characters with 200 overlap, per project or per file
- Sources
- Top 5 matches per search by default, adjustable up to 20
- Answer cache
- Exact for 1 hour, similar for 24 hours at 0.90
- Agentic loop
- Up to 5 sub-queries and 2 retrieval rounds
The API
Each project gets its own reference: 12 endpoints on one base URL, rate limits per key and per project, webhooks, and one command that connects Claude Code over MCP. SDKs for Python and JavaScript are on PyPI and npm.
-
API key
-
Install an SDK
-
Rate limits
-
Webhooks