Cyborgs Writing

Cyborgs Writing

Writing with AI

How do I organize my notes so AI can actually use them?

Getting started with your markdown AI database

Lance Cummings's avatar
Lance Cummings
Oct 09, 2026
∙ Paid
Image generated by gamma.ai

I experimented with a search of my own knowledge base before teaching a workshop on building one. It swept 999 files, opened 21, and missed every folder where the answer actually lived. Here is what changed on the second run, and the six-activity sequence I ran in Kraków to get there.


I spent this summer building something I could use to explore how structured content works with AI at scale ... while also providing my AI tools with a more nuanced knowledge base that is easy to access, but private and secure.

This meant building a markdown database on my hard drive.

If you don’t know what markdown files are, they are simply text files with simple notations that indicate formatting and content types.

Giving AI everything you've written is not the same as giving it something it can use.

In September I took it to the CAKE conference in Kraków and ran a ninety-minute workshop where people built one from scratch.

Before that, I wanted to test how AI actually works with this kind of knowledge base.

It’s been pretty clear that structure changes retrieval. How you organize files determines what an AI finds when it goes looking.

Easy to assert. Harder to show, because retrieval happens out of sight.

You see the draft, not the path to it. So I built a way to watch the path: a scan of my own vault that logs every thing AI does as it searches the database.

The test task was writing an AI policy for a new class ... something I write and revise every semester (and something my knowledge base ought to be good at).

Tracing how AI uses my knowledge

I have written those policies before, taught with them, and published about them.

The first run had the whole vault without any guiding documentation (no AGENTS.md, no domain map). It swept 999 files and opened 21. Not one of the 21 came from pedagogy or Teaching, which is where much of my information on this topic lives.

What it opened instead was adjacent material. Newsletter posts about AI and writing. A conference note. Documents that share vocabulary with an AI policy without being one.

The draft that came back was okay, but not quite right.

The second run had the same vault, the same model, the same prompt, with governing documents added:

  • a file telling the assistant what to read first,

  • a map of what each folder holds and who it’s for, and

  • specific skills for how to use the knowlege.

It swept 349 files and opened 10. Six of them were actual AI policies.

The counts are the easy part to report. What matters is what came out.

Here is the policy the first run wrote, trimmed but not otherwise touched:

Writing has always happened in a network between people and technology, and AI is now part of that network. In this course we will treat AI as a tool that can augment your thinking rather than replace it.
← ai-and-writing/ (public posts)

Students bring different mindsets to these tools — instrumentalist, formalist, humanist, technologist, posthuman — and part of our work this semester is recognizing which one you hold.
← My Students Are Mostly Instrumentalists About AI.md

AI systems retrieve fragments rather than whole documents, so the structure of what you give them shapes what you get back. Consider what kind of information you are asking for — a concept, a reference, a process — and structure your prompt accordingly.
← AI System/Information Type Decision Guide.md

Read it as a student. It is about AI and writing. It tells you nothing you can act on: no permission, no limit, no line you could cross.

That third paragraph is the giveaway ... it’s retrieval methodology from my own system folder, dropped into a syllabus, where it answers a question no student asked.

Here is the second run:

AI can both hinder and enhance our capacity to learn. We need to be mindful of when it impedes us and when it gives us new understanding. This class is a safe space to experiment with AI without shame, fear, or guilt.
← A Heuristic for Writing Your AI Syllabus Language.md

You will not be penalized just for using AI in this course. Unless I say otherwise for a specific assignment, use AI writing assistants freely.
← My AI Policy Wide Open With Guardrails.md

Generated text — permitted, up to 150 words per instance. You are the head writer. AI assistance — permitted without limit: ideas, outlines, feedback, audience adaptation.
← Give Your Syllabus AI Toggle Switches.md

I reserve the right to ask you if or how you used AI — not to get you in trouble, but so we both understand how the tech is shaping your writing and my teaching.
← What Students Asked After I Rolled Out My AI Policy.md

Same question, same assistant, same machine. The second one is a policy. It has a number in it.

It also has the sentence I rewrote six months after the first version, because students kept asking what the disclosure was actually for.

The run found the revision instead of the original, which was impressive.

We spend a lot of time talking about hallucination, and for this kind of work I think it’s the wrong word. Nothing was invented, really.

The assistant inferred a policy from content that was never meant to answer the question. The answer was still a good possible answer, but not customized to my work in the best way.

Worst of all, the AI had no clue. So if I didn’t know what was in that knowledge base, I would totally miss it.

The trend right now is to give AI all you got. Connect your notes. Sync your drive. Upload the folder. And AI will magically deliver wisdom and good answers.

But I started building this way for the opposite reason.

I wanted control over what AI could reach, which is most of why I separate storage, structure, and access in the first place.

What the scan showed me is that control isn’t only about what AI can reach. Something in the folder has to tell it what it is looking at.

The original LLM wiki … and half the story

In June, Andrej Karpathy posted about how he’s been spending his tokens lately — less on code, more on knowledge.

He points an LLM at a folder of raw sources and has it compile a wiki: markdown files, backlinks, articles for the concepts it finds.

He reads the result in Obsidian. He queries against it. He files the answers back in, so his research accumulates instead of disappearing into chat history.

Two things in that post stayed with me.

The first is what he didn’t need. He went in expecting to reach for retrieval infrastructure: vector databases, embeddings, the apparatus that has grown up around getting AI to use your documents.

He found that index files and short summaries of each document were enough, at least at the scale of a hundred articles. Markdown in folders, with a few files describing what’s in the other files.

That is the same finding I got from the second scan. I was measuring why a search went wrong. He was noticing that a simple setup went right. Both point at the index or guiding documentation doing the work.

The second is where we part company. Karpathy says of the wiki, “I rarely touch it directly.” The LLM writes it and maintains it; his hands stay on the raw sources and the questions.

For what he’s doing, that makes sense. He’s compiling research on topics he wants to understand, and the wiki is scaffolding for understanding them.

The domains in my vault are not a filing convenience. They're how I understand my own work.

My knowledge base is my teaching, my published writing, the arguments I’ve been making for years and the notes underneath them.

If an assistant reorganizes it according to categories it derived, I’ve handed over my expertise.

The domains in my vault are not a filing convenience. They’re how I understand my own work.

So I keep the taxonomy, but let the llm do the grunt work. The assistant converts, files, tags, checks, logs, and drafts. It doesn’t decide what the categories are, what counts as a principle rather than a process, or which folder a note belongs in when the answer is unclear.

Those are the decisions that make the knowledge base mine, and they’re also the decisions that determine what gets found later.

I’m still digging into this approach, but if you want to join the exploration, below are the instructions from my workshop that help you start from zero.

Paid subscribers get access to my live public skills library, along with all my developing materials on building structured knowledge bases.

Need something more foundational? Check out my course on Writing with Machines. Still free for paid subscribers, because I haven’t gotten around to changing that!

Keep reading with a 7-day free trial

Subscribe to Cyborgs Writing to keep reading this post and get 7 days of free access to the full post archives.

Already a paid subscriber? Sign in
© 2026 Lance Cummings · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture