The AI Knowledge Base That Writes With Your Voice

An AI that writes in your voice isn’t prompted into existence. It’s built – slowly, deliberately, in files you maintain over months. The system prompt is what the model carries into every conversation. Most solo founders never set theirs up. The ones who do treat it like a junk drawer instead of a curated archive, and wonder why the output still sounds generic.

The knowledge base is the part of an AI workflow that compounds. The prompt is the part that doesn’t. A good system prompt gets you through one conversation. A good knowledge base gets you through the next 200. The trade-off, both ways: setup cost is real, maintenance cost is real, and the output gets steadily more recognisable as yours the longer you keep at it. The setup is small. The discipline is what makes it work.

Three file types, three purposes

A working knowledge base has three file types. They live in three folders, three custom-instruction blocks, or three project knowledge sets – the structure matters less than the separation. Mixing them produces the same kind of context drift that the AI plateau article describes at the conversation level. The separation prevents it at the file level.

Voice files carry how you sound on the page. Sentence rhythm. Specific phrases you actually use. The constructions you avoid. Examples of paragraphs you’ve published that you’d publish again. Voice files are the smallest in count and the most important by weight. They’re also the hardest to build because most writers can describe their voice in adjectives but can’t give the AI usable examples without going back into the archive and choosing on purpose.

Positioning files carry what you’re for. The reader you write to. The position you take against alternatives. The topics you’ll write about and the topics you’ll never write about. Positioning files are the second category in importance. They prevent the model from drifting toward generic angles because they encode what makes your work distinct.

Reference files carry everything else. Source material you draw on. Past articles for context. Frameworks you’ve built. The bibliography of ideas you reuse. Reference files are the largest in count but the least load-bearing per file. They make the model’s output factually accurate without making it sound like you.

The split matters because each file type has a different retention rule and a different review cadence. Voice files almost never retire – they encode something stable. Positioning files retire when your position shifts. Reference files retire constantly as the underlying material ages out or becomes irrelevant.

Setting up the voice files

Voice files are where most knowledge bases fail at the start. The instinct is to write a description: “I write in short, direct sentences with occasional longer reasoning paragraphs. I avoid hype.” The model reads the description and produces the average of writers who describe their voice that way. Which is not you.

The version that works uses examples, not descriptions. Three to five files of 800–1,500 words each, each one a piece of published work that represents the voice you want to keep. The model learns from the examples in a way it can’t learn from the description. You’re showing, not telling.

Picking the right examples matters. The pieces should be:

  • Recent enough to represent current voice, not voice from three years ago
  • Varied across topics, so the voice signal separates from the topic signal
  • Things you’d genuinely publish again if you wrote them fresh today

Add one short file with notes about what makes those examples work – the constructions to keep, the constructions to avoid. The notes file is the description-style content. Its job is to reinforce the example files, not to replace them. Three to five examples plus one notes file is the smallest version of the voice layer that holds.

Setting up the positioning files

Positioning files are about your work, not your taste. They answer three questions:

  • Who do you write to, specifically? Not “solo founders” but the version of a solo founder that maps to your actual reader – age range, business model, what they already know, what they’re stuck on.
  • What position do you take against alternatives? The articles, books, frameworks you’re explicitly arguing with. The angle you’re not taking. The audience you’re not writing to.
  • What topics are in scope, and what topics are out? The intersection of what you write about and what you refuse to write about is where positioning lives.

One file per question is enough, usually 300–800 words each. The positioning files don’t need examples – they encode editorial decisions, not voice, so descriptive prose works fine.

The positioning layer is the one that catches the most early drift. Without it, the model defaults to writing for the broadest plausible audience, which is the average audience, which is no audience in particular. With it, the model writes for the reader you actually have.

Setting up the reference files

Reference files are the easiest layer to build and the easiest to mess up by accreting too much into.

What belongs:

  • Past articles you reference repeatedly (linkable canonical statements of your frameworks)
  • Source material you draw on (research papers, books you cite often, primary documents)
  • Frameworks you’ve built (the named structures that recur across your work)
  • Domain knowledge specific to your business (services, products, prices, processes if you reference them in client-facing writing)

What doesn’t belong:

  • Random research notes you might use someday
  • Articles by other people you found interesting once
  • Background reading that doesn’t surface in your actual writing
  • Old drafts of things you never published

The trap most solo founders fall into: treating the reference layer as an inbox rather than an archive. The 2023 paper Lost in the Middle showed that language models pay less attention to information in the middle of long contexts than at the beginning or end. The implication for a knowledge base: more files isn’t better. Past a certain size, the reference layer starts diluting the voice and positioning layers, because the model is splitting attention across material that isn’t load-bearing.

A clean reference layer holds 10–25 files most of the time. If yours has 60+, half of them are reference material that should have stayed in your reading notes and never made it into the AI’s context. The fix is the maintenance routine, which is the second half of the system.

The maintenance routine

A knowledge base that gets built and never pruned drifts in exactly the way an unmaintained codebase drifts. The discipline is small and runs weekly.

Andy Matuschak’s writing on evergreen notes makes the case for curation as the active part of a personal knowledge system. The same principle applies to an AI knowledge base. The work isn’t adding files. It’s deciding what stays.

The routine:

  • Every week, one pass through the reference layer. Ask of each file: did I draw on this in the last 30 days? If yes, keep. If no, mark for retirement.
  • Every quarter, one pass through the positioning files. Has my actual writing diverged from what these files claim? If yes, update or retire.
  • Every six months, one pass through the voice files. Do the examples still represent how I want to sound? If a voice file is producing output that no longer fits, retire it and pick a fresher example.

The retirement test is the load-bearing part. A file retires when the output it produces no longer sounds like work you’d publish today. Not when the file is old. Not when the file is wrong in some technical sense. When the output drifts away from yours, the file is the cause more often than the prompt is.

This is the move most knowledge bases skip: when output sounds off, the instinct is to refine the prompt. The right move is usually to retire a file. The prompt is what you say to the model. The knowledge base is the lens it sees through. A blurry lens produces blurry output regardless of how clear your prompt is.

When the whole base needs a rebuild

Rare, but real. Roughly once every 18 to 24 months, the entire knowledge base accumulates enough drift that incremental retirement stops keeping up. Three signs:

  • The output sounds generic across multiple conversation threads even after starting fresh with the plateau-crossing patterns.
  • Multiple files are producing output you no longer recognise as yours.
  • You can’t remember the last time a knowledge-base entry made the output noticeably better.

When two of the three are true, rebuilding is faster than refining. The rebuild looks like this: start with a fresh, empty base. Add three voice files using your most recent representative work. Add three positioning files for the current version of your position. Add five reference files for the material you’re actually drawing on this quarter. Run the rebuilt base for 30 days. If the output is closer to your voice than the old base produced, keep going. If not, the issue isn’t the base – it’s upstream, in the editorial decisions you’re making about what to write.

The rebuild is uncomfortable because it discards 18 months of accumulated files. The cost feels real. The actual loss is small. The bulk of what gets discarded was already producing the drift the new base is meant to fix.

A note on tools

The knowledge base is not the tool. The tool is whichever AI you happen to use this year. The files are the system that survives the tool change.

The same files work across Claude Projects, ChatGPT Custom GPTs, Gemini Gems, and whatever launches next. The format is text. The structure is three folders. The principle is curation over accretion. Solo founders who get attached to a specific tool’s knowledge-base feature spend energy migrating each time they switch. Solo founders who treat the files as the system spend a Sunday afternoon copying files into the new tool’s interface, and the new tool starts producing in their voice on the first conversation.

The rest of the AI & Workflows archive – including the cluster closer on when to ignore AI output entirely – lives at the AI & Workflows category.

Curate the files

The knowledge base is the part of an AI workflow that produces the output you’d describe as yours. It compounds in a way prompts don’t. It also drifts in a way prompts don’t, which is why the maintenance routine is the load-bearing half of the system.

Three file types. Weekly pruning. Quarterly review. A rebuild every 18 months if the drift gets ahead of the maintenance.

Curate the files. The voice follows.


Build Your Content Machine. A free 5-part email course on building an AI content system that sounds like you – the knowledge base above is one layer of it. Start the free course →

Did you like this article? Share it with a friend!