Project 03

LLM Applications and Agents

Learning to build applications on large language models, in two topics: the harness, the part outside the model that lets it carry out tasks as an agent, and RAG, which has the model answer from external sources. Notes are filed by topic, and each starts from a concrete question.

Status
Planning
Started
September 2026
Planned tools
Python · FAISS · pgvector · MCP · LangGraph · RAGAS

This project has not started yet. For now, this page shows the plan and the structure future entries will follow.

What I want to learn

Understand what problem each part of an agent and of RAG solves and how it works underneath, and check the key parts by writing them myself.

Development plan

  1. 01
    In progress

    The agent loop and a minimal RAG

    Write an agent loop without a framework, then plug a minimal RAG with citations into it as a tool.

  2. 02
    Planned

    Inside vector search

    How IVF, PQ and HNSW work and how their parameters trade recall for speed, measured with FAISS, then a simplified index written by hand.

  3. 03
    Planned

    Context engineering and retrieval quality

    Manage the context budget and caching; compare BM25, dense and hybrid retrieval with reranking.

  4. 04
    Planned

    Tools, MCP, skills and subagents

    Build an MCP server, a skill that loads on demand and a read-only subagent.

  5. 05
    Planned

    Hooks, sandboxes and permissions

    Block dangerous actions with a hook, run code in a sandbox, limit tools and skills by role and filter retrieval by permission.

  6. 06
    Planned

    Memory, durable execution, ingestion and generation

    Build per-user memory and resumable execution; compare chunking strategies and check citations.

  7. 07
    Planned

    Evaluation, observability and production

    Evaluate agents with pass@k and pass^k, export traces and make the RAG multi-tenant.

Topics

Notes are filed by topic; each new note goes under the topic it belongs to.

01

RAG

RAG: retrieval-augmented generation

Having the model answer from external sources: ingestion and indexing, how vector search works, retrieval quality, generation, evaluation and running it in production.

0 entriesView notes
02

Harness

Harness: everything around the model that makes an agent

Everything outside the model: the agent loop, context engineering, tools and MCP, skills, hooks, sandboxes, permissions, memory, durable execution, evaluation and observability.

0 entriesView notes

Overall study roadmap

The full study path for both topics, Harness and RAG: the order to learn them in, how deep to go on each point, common questions and sources.

Read the roadmap

Entries

As the project develops, I will add the questions, implementation notes, and test results from each stage here.