(CASE STUDY · 02)

OREAG

  • RAG
  • Agent Memory
  • MCP
  • pgvector
  • FastAPI · Next.js

Oreag is a RAG and memory service. A developer uploads documents, tunes chunking, the embedding model and the answer model with their own provider keys, and gets a per-project API that answers with the file, page and similarity of every source. Coding agents use the same project as memory, through an MCP server.

ROLE
Product & Engineering
TIMELINE
7 Months
YEAR
2026
TEAM
Solo

(MY ROLE)

  • Designed and built Oreag alone: the Next.js dashboard, the FastAPI backend, the pgvector schema and the MCP server.
  • Built a read path that searches two ways and asks back when the evidence is thin.
  • Built usage metering, a query log and a Gaps inbox that turns flagged questions into evaluator tests.
  • Wrote the in-app docs, a C4 model and a CI check that fails when either drifts from the code.
(WHY IT EXISTS)
An app needs answers from its own documents. A coding agent needs a place to remember. Neither should mean running another vector database, so one Oreag project serves both.
(THE FIRST VERSION WAS A SCRIPT)
February 2026: a LangChain script. Text files in, 1000-character chunks with 200 overlap, OpenAI embeddings, a Chroma collection. The chunk defaults survived. Chroma did not.
(WHAT CHANGED)
A deploy hook failed every file in flight, so the files table became a queue with leases. Keyword search matched strings, not words, so it now stems. The semantic cache matched topics, so its floor rose from 0.75 to 0.90.

Process

  1. 01

    Prototype & Research

    A 36-line LangChain indexer in February, a LangGraph reference in April. On June 11 the script moved to legacy/ and the platform build began.

  2. 02

    Platform & Keys

    The dashboard and API landed on June 12, bring-your-own keys on June 17, encrypted with Fernet. Each project got its own /v1 endpoint and oreag_sk_ key.

  3. 03

    Memory & Agents

    Agent memory landed on June 18: a memories table, an MCP server and a Memory tab. Memories shared the chunks' vector space, and the 3D graph followed on July 3.

  4. 04

    Retrieval & Trust

    Hybrid search and the semantic cache landed on July 5, stemming on August 13. September added the evaluator, Queries, Health and Gaps, to find where answers fall short.

How It Works

Upload a file and a worker converts, chunks and embeds it into pgvector. Ask a question and the API checks its cache, searches two ways, and answers only when the evidence holds. Agents reach the same project through MCP.

The Workspace

Every project is one page with six tabs: Files, Memory, Playground, API, Visualize and Settings. Files shows what is indexed and which edition is in force. The Playground runs the same code path as the public API, so an answer here is an answer there. Both screens show sample data from Oreag's own docs.

In The Detail

Every answer carries its sources: file, page and match score, with the cited ones marked. Agents keep notes in the same project, and pinned ones come first in the list an agent reads when a session starts.

Under The Hood

Bring your own key for any of 16 hosted providers, or run one of 3 local runtimes with no key. Keys are encrypted at rest and resolved per request: the project's own key first, then the account's. A new project starts from the defaults below.

(SCOPE)

file extensions on the upload allowlist
29
model providers, keyed or local
19
MCP tools for coding agents
9

THE DEFAULTS

Chunking
1000 characters with 200 overlap, per project or per file
Sources
Top 5 matches per search by default, adjustable up to 20
Answer cache
Exact for 1 hour, similar for 24 hours at 0.90
Agentic loop
Up to 5 sub-queries and 2 retrieval rounds

The API

Each project gets its own reference: 12 endpoints on one base URL, rate limits per key and per project, webhooks, and one command that connects Claude Code over MCP. SDKs for Python and JavaScript are on PyPI and npm.

  1. API key

  2. Install an SDK

  3. Rate limits

  4. Webhooks