# Voice Compiler

Canonical: https://andydataguy.com/wiki/ai-and-technical-development/default/voice-compiler

Author: Anand Houston (AndyDataGuy)

One piece of content goes out on three channels. The cloth stays the same, each cut fits the people reading on that channel, and one wall of shared minimums stands behind all three.

Many job posts I've read from people who buy writing at volume describe the same week. A draft comes back sounding like nobody in particular. The buyer fixes a sentence, and the next draft carries the same sentence again. By the third time they're fixing the same sentence, they're paying for a writer and doing the writing.

The owner of an equipment dealership told me on a call about a batch of social posts that he can't judge grammar and wasn't going to try. He sells construction positioning gear to the people who run excavators and graders, and the posts were for them. He said he'd read each one to make sure it made sense and that the posts "have the right vibe and they sound right." I'd worried out loud that one product description might be talking nonsense about how the blade works. He read it and said, "It's talking correctly."

He'd split the job in two.

The half he handed off, grammar, is work that written rules and a careful reader can carry. Getting his posts there took more than a dozen rounds of drafts, edits and reads. On one round, a reviewer model that hadn't written a word of them graded 33 rewritten captions against the written rules and the owner's own recorded words, and flagged 10 as major problems. After the next revision it flagged none.

The half he kept is whether the posts make sense to an operator and talk like someone he'd trust. I can't write a rule that decides that. His ear decides it, or the ear of someone who sits in a cab all day.

Those halves are two different machines. The first decides what a piece says. It breaks the piece into parts, gives each part a job, and cuts whatever doesn't earn its place. I call it a content compiler. The second takes content that already earns its place and makes it sound like one business talking to its own customers on one channel. I call that one a voice compiler. Both stand on a short list of minimums every sentence has to meet, whoever the client is: no em dashes, no claim without someone who saw it, no headline that stops before its thought does.

I've been running both by hand on client work, with models drafting and a person or a second model checking every operation, to find out whether the quality is reachable before anyone builds the software. Some of that quality is reachable now. The rest needs the client's ear, and my method treats that ear as data to collect.

Here's the test the split has to pass. Say an agency builds websites for ten chiropractors. The clinics share a service list and most of a page structure, so most of the content can be shared. The voice can't. If two of those sites read alike, the voice compiler did nothing, and a patient comparing two clinics in the same town will hear it.

## AI copy sounds like everybody for three separate reasons

Three stacked bands. Content asks whether each part does a job for the reader. Voice asks whose mouth the words come out of. The tells, at the base, are habits readers spot, and grammar is a thin line inside that base band.

THREE WAYS A DRAFT SOUNDS LIKE AI

CONTENT

Does each part do a job for the reader?

An intro that restates the headline.

VOICE

Whose mouth is it coming out of?

Fine sentences that could run on any competitor's page.

THE TELLS

Would a reader spot the habits?

Em dashes, inflated verbs, an invented quote.

GRAMMAR

Fixing one leaves the other two where they were.

When a draft sounds like AI, three separate things can be wrong: parts with no job, a voice that belongs to nobody, and the tells. Grammar is one thin line inside the smallest of them.

A language model writes toward the average. Ask it for a post about grading blades and it returns the middle of every post about grading blades it ever absorbed, polished and balanced and safe. The writing standard I hold my own drafts to defines bad machine writing as "average generalization on a concept, delivered in a way that helps nobody learn." You get the cadence of a consultant, the vocabulary of a conference talk, and no particular person underneath.

When a buyer says a draft "sounds like AI," three separate failures are hiding under that phrase, and each needs its own fix.

The first failure is content. Parts of the piece do no job for the reader: an intro that restates the headline, a paragraph explaining a problem the reader already lives with, a summary that drops the condition the body was careful to state. The second is voice. Every sentence is fine and none of them belong to anyone, so the post could run on any competitor's page without a word changed. The third is the tells, the surface habits readers have learned to spot: em dashes, the same dozen inflated verbs, the "it's not X, it's Y" turn, the line that informs you what you're feeling.

Fixing one leaves the other two where they were. I've seen drafts scrubbed of every em dash and every inflated verb that still had no point. I've seen drafts with a sharp point that read like a brochure. A banned-word list can only ever reach the third failure.

One summer draft for the dealer shows all three at once. It opened, "Ask a survey lead how the first morning of the season went and the best answer is a quiet one." It went on to quote a crew member nobody had interviewed, and it closed on a tidy line about plain talk beating polish. The content failed first: the post never said what a reader gets, and the owner sent it back saying he couldn't tell what it was about. The voice failed with it. My own note on drafts like that one was "Who is the human? Speak to them. This is not speaking to anybody." The tells were the invented quote and the aphorism for a closer. Cleaning up the tells alone would have left a grammatical post about nothing.

In the one place I've measured it carefully, the tells were also the smallest part. I had fifteen articles I'd written for an AI-tools publisher reviewed line by line by a separate model working from a written editorial standard, and the client had already approved every one of them. The review logged 1,368 observations. The biggest category was claim logic and evidence fit, with 334 notes spread across all fifteen articles. Technical and product precision came next with 297, then sources and attribution with 217. Grammar drew 28 notes. Mechanics, the commas and capitals and hyphens, drew 30.

Bars for review notes by category across fifteen approved articles: claim logic and evidence fit 334, technical and product precision 297, sources and attribution 217, word choice 100, voice and stance 74, concision 74, mechanics 30, grammar 28.

1,368 REVIEW NOTES ON 15 CLIENT-APPROVED ARTICLES

Claim logic and evidence fit THE CLAIMS

334

Technical and product precision THE CLAIMS

297

Sources and attribution THE CLAIMS

217

Word choice and terminology

100

Voice and stance

74

Concision and repeated meaning

74

Mechanics THE COMMAS

30

Grammar THE COMMAS

28

THE CLAIMS

THE COMMAS

Selected categories. These count review notes, not error rates: one problem can earn several notes.

A line-by-line review of fifteen client-approved articles logged 1,368 notes. Claims drew about twelve notes for every one about grammar, so the problems sat mostly in the claims and their evidence.

The reviewer was careful to say those are counts of review notes and not error rates, since one problem can surface in a summary, a table and an FAQ and earn three notes. Even allowing for that, claims drew about twelve notes for every one about grammar. The reviewer's summary put it plainly: "Grammar deserves attention, but it is not the best organizing explanation for the work."

So I tried having a second model check the first one's work, and measured what that bought. Three drafts had passed every string check I'd built at the time: no em dashes, no banned words, clean on every lint rule. Two model reviewers, each blind to the other, scored them 77 and 88 out of 100. I scored the same three around 30 and called them obviously machine-written.

My read on the 58 points between the higher score and mine is that they come from where the judge learned what good writing looks like. It learned from the same pool the writer did, so the writer's habits read to it as quality. A model checking a model is mostly checking whether the text is likely, and likely text is what the writer was built to produce. Two judges trained on the same pool agree with each other for the same reason, which is why two of them didn't help.

So the work splits the way the failures split. Whether each part earns its place gets its own check. Whose voice it's in gets another. Underneath both sits a set of minimums a pattern or a careful reader can enforce, and at the end there's a human ear that no model stands in for. Editors were dividing their work along those lines long before anyone had a model to worry about.

## Editors already split this work

A manuscript at a publishing house goes through several readers, and each one reads a different unit of the text. One reads the whole book and asks whether chapter four should exist. Another works paragraph by paragraph on how the prose moves. A third fixes commas, spelling and cross-references, and a last one checks the typeset pages. The comma reader isn't the one who tells the author the middle third of the book is in the wrong order.

Editors Canada writes the first of those jobs into its professional standards. [Structural editing](https://www.editors.ca/publications/professional-editorial-standards/structural-editing/), in its words, "is assessing and shaping the overall organization and content of the material to optimize it for the intended audience, medium and purpose." The standard tells the editor to make deletions "to remove repetitive, irrelevant or otherwise superfluous text or other elements" and additions "to fill gaps in content or strengthen transitions between sections."

Four rungs of a ladder. A structural edit reads the whole piece and maps to the content compiler. A line edit reads paragraphs and sentences and maps to the voice compiler. A copyedit maps to the shared minimums. A proofread maps to a check of the live page.

WHAT EACH READER READS, AND WHAT IT BECOMES

STRUCTURAL EDIT

Reads the whole piece and its parts

CONTENT COMPILER

LINE EDIT

Reads paragraphs and sentences

VOICE COMPILER

plus a written voice profile per client and channel

COPYEDIT

Reads sentences and words

SHARED MINIMUMS

PROOFREAD

Reads the final page

LIVE-PAGE CHECK

after it ships

Publishing houses already split editing by the unit each reader works on, from the whole book down to the final page. Each rung maps to one machine, so the old split can run on every piece.

The Editorial Freelancers Association [defines the levels below it](https://www.the-efa.org/editorial-services-definitions/): line editors work "at the sentence or paragraph level," copyediting means "correcting spelling, grammar, usage, and punctuation," and proofreaders check the finished pages.

When I describe what I want from a content system, I end up asking a structural editor's questions. Does this scene add value? Does this character have the impact it's looking for? Is the overall arc worth making? A content compiler puts those questions to a blog post's third section or a landing page's proof block, every part of every piece, and writes each answer down.

The voice compiler maps to the second rung, the stylistic edit, with one addition: a written profile of how this client talks, per channel. The copyedit maps to the shared minimums, and the proofread to checking the live page after it ships.

Voice needs a definition before it can be specified. One email-marketing company's public style guide draws the line I use: "[Our voice doesn't change much from day to day, but our tone changes all the time.](https://styleguide.mailchimp.com/voice-and-tone/)" Voice is who's talking. Tone is how that person adjusts to the moment, the way you'd talk differently to a customer whose order went missing and one who just placed a second order.

Nielsen Norman Group breaks tone into [four dimensions](https://www.nngroup.com/articles/tone-of-voice-dimensions/): formal or casual, serious or funny, respectful or irreverent, matter-of-fact or enthusiastic. Those dials describe a difference between two pieces of copy. They're too coarse to produce one: two chiropractic clinics can sit at the same setting on all four and still sound identical, because the dials say nothing about the words a clinic uses, the proof it trusts, the jokes it would never make, or which patients it's talking to.

Voice is who's talking, and it stays put. Tone is how that person adjusts to the moment, quieter for a customer whose order went missing and brighter for one who just ordered again.

The sociolinguist Allan Bell studied that last item, how speakers shift their style for whoever's listening, and his finding opens the paper: "[Style is essentially speakers' response to their audience.](https://www.cambridge.org/core/journals/language-in-society/article/abs/language-style-as-audience-design/35677FBB8C9B7602DC20FDD354DB2ADD)" People adjust mostly toward whoever they're addressing. In mass media, Bell found, speakers also shift toward an absent group they identify with, a group he called referees.

Bell's referee idea explains why a dealer's posts should sound like the operators it serves. The posts are read in a cab by people who run machines, and the voice answers to them, even though the dealer is the one paying for the writing. Bell's finding also sets the unit a voice compiler works on: the client, talking to one audience, on one channel. Change any of the three and the voice has a different job.

None of these passes is new. Running all of them on every part of every piece, for many clients at once, without the structural questions getting skipped when the calendar fills up, is the part that needs a machine, and the machine starts with the content side.

## The content compiler now judges what each part is worth

My earlier wiki entry on the [content compiler](/wiki/ai-and-technical-development/default/content-compiler) laid out the architecture, and the architecture holds. A piece of content is a tree of typed parts instead of one long string: a hook, a section, a claim, a citation, a call to action. In that design, each part is produced and checked on its own, so a revision to one paragraph re-renders that paragraph and serves the rest from cache. One content tree can be emitted to many channels the way a compiler emits one program to many machines. The [reference build](/case-studies/content-compiler-reference) shows that architecture in practice: each paragraph is stored as its own row, and a paragraph that fails a check gets regenerated on its own while every approved paragraph stays exactly as it was.

That entry never said what the checks on each part are for. Its passes confirmed that a part existed, that it was formatted, that its citation resolved, that it obeyed the style rules. A paragraph can clear every one of those checks and still do nothing for the reader. Running the operations by hand on client work showed me what a pass is for: deciding whether each part earns its place.

So the content compiler, as I run it now, starts by breaking a piece into the parts its type calls for. A website breaks into pages, each page into sections, and each section into elements: a headline, a label, a button, a caption, a line of alt text. An article breaks into a headline and deck, an opening, body sections that each carry a claim with its evidence and example, a close, and the short forms a skimmer reads (summary, FAQ answers, takeaways, the search description). A social post breaks into four parts: the hook, the point, the proof and the ask.

Three trees of parts. A website splits into pages, sections and elements. An article splits into headline and deck, opening, body sections, close and short forms. A social post splits into hook, point, proof and ask. Each part carries a one-line job.

WEBSITE

- Pages : one per visitor question, such as a service page Sections , such as the proof block: makes the claim above it believable to someone who's never heard of the business Elements : headline, label, button, caption, alt text

ARTICLE

- Headline and deck : one complete thought a stranger could repeat
- Opening : puts the reader inside the problem before the method arrives
- Body sections : each carries a claim with its evidence and example
- Close : tells the reader what to do next
- Short forms (summary, FAQ answers, takeaways, search description): say nothing the body doesn't already say

SOCIAL POST

- Hook : names the situation the reader is in
- Point : one benefit the viewer gets
- Proof : the one sourced fact that makes the point believable
- Ask : what the viewer should do next

A part whose job can't be stated in a sentence is a cut candidate.

Every content type breaks into parts, and every part gets a job written as what it does for this reader at this point. A part whose job you can't state in a sentence is the first thing to cut.

Each part then gets a job, written as what it does for this reader at this point. A service page's proof block exists to make the claim above it believable to someone who's never heard of the business. An article's opening exists to put the reader inside the problem before the method arrives. A part whose job can't be stated in a sentence is a cut candidate, and the structural editor's moves follow from there: cut what's repetitive or irrelevant, add what fills a gap, move what's out of order, fix what does its job badly.

Each part also carries a rule about how much it may make up, and the rule depends on the part's type. A part that asserts something about the world traces to a source, or it stays a visible gap in the draft until someone fills it. A part that condenses the body, like a summary or an FAQ answer, may say nothing the body doesn't already say. A part the writer composes, like an analogy, a hook or a transition, has to teach or move the reader along, and any fact tucked inside it follows the sourcing rule. My standard puts the line in one sentence: you may invent a metaphor, and you may never invent a story, a figure, a quote, a source or a specific detail.

What comes out of these steps is a content map: every part with its type, its job, a verdict (keep, cut, add, move, or a question for the author) and the evidence behind any claim it makes. The map is what the voice compiler receives. It never receives a loose draft.

Order is often a content decision. An audit of one batch of the dealer's queued posts found that 23 of 34 distinct first lines answered a question before the post had raised it. One draft post for the dealer opened, "You can get help on a machine without anyone driving out to the job." It's a clear sentence, and it's the answer to a question the reader hasn't been asked yet. My rewrite opened on the situation: "A machine failing outside of town can mean hours or days before support arrives." The old opener survived as the second line, almost word for word. Nothing about the wording was wrong. The parts were in the wrong order, and the first part had no job until the trouble had been named.

Before: a draft post opens with "You can get help on a machine without anyone driving out to the job," stamped move. After: a new first line names the trouble, stamped add, and the old opener follows as the second line, almost word for word.

ONE POST, JUDGED PART BY PART

BEFORE

MOVE You can get help on a machine without anyone driving out to the job.

Answers a question the reader hasn't been asked yet.

AFTER

ADD A machine failing outside of town can mean hours or days before support arrives.

Names the trouble first.

MOVED You can get help without anyone having to drive out.

The old opener, almost word for word, now answers the line above.

VERDICTS KEEP CUT ADD MOVE ASK THE AUTHOR

A draft post answered a question before the reader knew there was a problem. The rewrite opens on a machine failing outside town and keeps the old opener, almost word for word, as its second line.

Next comes a content check: claims against their evidence. A capability can't be stated as a guarantee. A single experience can't be stated as a law. A method that worked once can't become the best method. These checks run on content, before anyone touches the voice, because a sentence that claims too much is wrong in every voice, and polishing it only makes the overclaim more persuasive.

Then the short forms get read by themselves. Headline, summary, FAQ answers, captions, the call to action, the search description: read only those, then write down what a skimmer would take away, and compare it with what the body says. When the two disagree, the short form is wrong, because the body is where the writer had room to state the conditions.

One line from that fifteen-article review shows how a condition goes missing. The draft said, "Grounding does not make an AI incapable of being wrong. It makes it checkable." The reviewer's revision kept the first sentence and changed the second to "When the system exposes its sources, it makes the answer easier to check." The memorable version is the one that gets lifted into a takeaway or a social card, and it's the version missing the condition that makes it true. A skimmer who reads only the short form walks away believing grounding makes answers checkable by itself.

Two versions of one passage. The first ends "It makes it checkable," marked as the line a takeaway box would lift. The revision ends "When the system exposes its sources, it makes the answer easier to check," with the restored condition marked in amber.

THE MEMORABLE VERSION

Grounding does not make an AI incapable of being wrong. It makes it checkable.

WHAT A TAKEAWAY BOX LIFTS

THE REVISION

Grounding does not make an AI incapable of being wrong. When the system exposes its sources , it makes the answer easier to check.

THE CONDITION, PUT BACK

A skimmer who reads only the first version believes grounding makes answers checkable by itself.

The memorable line is the one that gets lifted into a takeaway or a social card, and it lost the condition that made it true. The revision keeps the claim and puts the condition back.

The social post is where all of this gets small enough to see at once. Before a word of a post gets written, it gets a four-line brief. The first line names the audience, who it's for. The second is the point, one sentence from the viewer's side saying what they should take away. The third is the proof, the one sourced fact that makes the point believable. The fourth is the ask, what the viewer should do next. Advertising has long had a name for the second line, the single-minded proposition: the one thing the viewer should walk away with, decided before any copy exists. My rule for that line is blunt: "If you can't answer that question, you have no business writing anything else."

The point has a second test, which I added after one draft passed the first. A point can be clear, sourced and correctly understood by a cold reader, and still be worthless to the viewer. One draft's point was that the dealer adds machine control to machines built without it. That's the seller's business model. No operator spends his evening worried about which of his machines shipped without machine control. The point has to be a benefit the viewer gets, like cutting to grade on the first pass, and a point about the company, its market position or its sales process fails before anyone judges a single word.

Every defect the content compiler finds goes back to the step that owns it, and gets fixed there. A structural hole goes back to the outline. A claim with no source goes back to sourcing, where it either finds a source or gets cut. A short form that drifted gets rewritten from the body's most careful sentence. The temptation with any finished-looking draft is to patch the sentence where the problem surfaced. I've watched that patch hold for one draft and then fail on the next, because the outline or the brief that produced the defect is still sitting there, instructing the next draft to make it again.

A defect shows up in a sentence and starts in the outline, the brief or the sourcing. Patch the sentence and the next draft drips again; fix the step that made it and the leak stops.

Run on a post, then, the content compiler answers four questions in order: who it's for, what they get, why they should believe it, and what they do next. Run on a website or an article, it answers the same kind of question about every part. None of those answers says anything yet about how the piece should sound. A post can pass all four and still read like it came from nowhere in particular, and that's the job the voice compiler picks up.

## Fifteen approved articles still had their biggest problems in the claims

The fifteen articles from that review were mine. I'd written them for an AI-tools publisher, and the client had approved every one before the review started. Approval meant the client was happy to publish. It didn't mean the articles were finished. The review read them line by line against a written editorial standard.

The reviewer was careful about what the evidence could carry. These were client-facing pieces, and some sentences carried the client's own edits or supplied marketing language, so the reviewer framed every pattern as something the corpus did, not as a verdict on how I write. I've kept that framing below because it's the right size for the claim.

The largest finding was the gap between a concrete experience and the general claim that came after it. The reviewer's own sentences: "A capability becomes a guarantee. A useful method becomes the best method." And the habit the reviewer recommended in its place: "ask what, exactly, the preceding evidence entitles the next sentence to say."

The before-and-after pairs make it concrete. One article told readers a connection would let an assistant "safely reach your real files, apps, and data." The revision dropped one word, "safely," because the sentence presented security as a built-in property of the connection. Another said, "You can add guardrails and evaluation loops to ensure quality control." The revision: "You can add checks and review passes to support quality control." The section after that sentence described review steps, so "ensure" promised more than the article went on to show.

Three before-and-after cards. "To ensure quality control" becomes "to support quality control." "It's imperative for people to always have alternatives" becomes "I realized I needed a backup." "Rising to become the best solution" becomes "That's where the product fits."

GUARANTEE TO CAPABILITY

You can add guardrails and evaluation loops to ensure quality control.

You can add checks and review passes to support quality control.

ONE STORY TO ONE PERSON'S LESSON

That's when I understood that it's imperative for people to always have alternatives to [the assistant].

That's when I realized I needed a backup to [the assistant].

BOAST TO JOB

rising to become the best solution

That's where [the product] fits.

Every fix sizes the claim to the evidence under it. None of them touched grammar.

Three fixes from a review of fifteen approved articles: a guarantee cut back to a capability, one person's story kept to one person, and a boast replaced by the product's job. None of them touched grammar.

Anecdotes got stretched the same way. One article told a personal story about relying on one AI assistant and concluded, "That's when I understood that it's imperative for people to always have alternatives to [the assistant]." The revision kept the story at its own size: "That's when I realized I needed a backup to [the assistant]." One person's experience became a rule for everybody in the first version, and a lesson about one workflow in the second. A line about a client project ended by promising something "you can reliably use for your business or studies," and the revision brought it back to "an answer I could use for the client."

Promotion arrived before proof, too. One article said the publisher's own product was "rising to become the best solution" in its category, and the revision cut the line down to "That's where [the product] fits." The shorter line gives the product its job in the reader's workflow and stops. A line saying prompt quality "directly determines" output quality became "helps shape," because a later section of the same article explained that model choice matters as well. The article had contradicted its own sole-cause claim a few scrolls down.

None of those fixes touched grammar. Every one is a content check, the kind the previous section described: does the claim match what the evidence underneath it shows?

The second finding was about voice. The review's strengths section argued for protecting things a cleanup pass would sand off. A concrete scene got a keep note: a sentence that began "I once worked with a heavy equipment client" and went on to name the exact equipment he sold. The reviewer's reason was that the detail establishes a real problem before any jargon appears. A candid setback got one: "The proposal didn't win that particular client." It keeps the later revenue figure from reading like a tidy success story, and the short sentence changes the pace. Two fragments got one too, "Evaluating. Refining.", because they slow the reader down at the practice the article is teaching, and what they're missing is obvious.

The reviewer then warned against the obvious fix: "Making every sentence formally complete and uniformly cautious would erase some of the writing's best qualities while leaving its conceptual problems untouched." A checker that flags every fragment, every long sentence and every strong claim produces a draft that's cleaner and worse. So the system I'm building counts what it keeps as well as what it flags, and a keep count that falls between drafts shows up as a warning that the checker is sanding off the voice.

A checker that flags every fragment, long sentence and strong claim leaves a draft cleaner and worse. This one cuts the overreaching claims, tags the scene, the setback and the fragment it kept, and counts them between drafts.

The review also set limits on the fixes themselves, and I adopted every one. A fix never invents an actor; when the doer is unknown, the fix becomes a question to the author. A fix never manufactures first person; the reviewer wrote, "Do not let an editor manufacture 'I tested,' 'I prefer,' or 'in my experience' to rescue an unsupported claim." A fix never repairs a bad number with another plausible number. And a fix never smothers a strong claim in hedges. The reviewer asked for the strong claims to stay "when they are warranted, rather than replacing confidence with a blanket layer of 'may,' 'might,' and 'can.'" Where a guarantee overreached, the replacement names the condition under which the claim holds.

Those limits matter most in the voice step, which is where a rewrite is most likely to change a fact while improving a sentence.

The review closed with five questions a finished article should let its reader answer: "What is the useful idea? What actually happens? Under what conditions is the claim true? What should I do next? Why should I trust this account?" Any buyer can put those five questions to any piece a vendor sends back.

Here's what I did with it. The review became six new checks in the rule list my checker reads, each with a plain name. One catches a capability stated as a guarantee. One catches a claim that quietly switches from one thing to another, such as from a model to the app built on it. One catches a summary that drops the body's condition. One catches an anecdote stretched into a claim about everybody. One catches a contrast drawn wider than the real distinction. One catches promotion that arrives before any proof. The keep notes became keep signals, which stop the checker from flagging a scene, a setback or a deliberate fragment.

The reviewer's advice on how to model my voice: "The most faithful model of your voice will be built from those distinctions, not from a list of forbidden words." The distinctions come from edits a person accepts or rejects, with the reason attached. A list of banned words can't hold that, and a record of decisions can.

## A voice compiler makes the same content speak for one client

The voice compiler receives a content map where every part already has a job and every claim already has its evidence. Its work is to make those parts sound like one business talking to one audience on one channel. Before it renders a word, it needs five inputs written down, because a voice nobody wrote down gets replaced by the model's default.

Who's speaking, the client's profile, the audience, the channel and a reference set get written down before drafting starts. Without those marks, every client's copy drifts to the same place: the model's default.

Input What it holds

Who's speaking Me in my own name, me under a client's byline, or the brand speaking as "we". Each one changes the pronouns, who can be quoted, and where a source may appear.

The client's voice profile The words this business uses and avoids, the kind of proof it trusts, the things it would never say, how often it contracts, and the usual shape of its sentences.

The audience Who reads it, what they already know, and the words they use for their own problems, collected from where they talk when nobody's selling to them.

The channel The constraints of the place it runs: length, how a post opens, who owns the page, what a reader is doing when it reaches them.

A reference set Finished pieces the client approved for this channel, used as examples to write toward and as the test the draft is scored against.

The channel input carries real numbers. The dealer's captions that I'd reviewed and kept had a median sentence of 12 words on LinkedIn, 10 on Facebook and 8 on Instagram. Contractions showed up at a median rate of one: in at least half the captions on each channel, every place a contraction could go, it went. Same company, same products, and three sentence shapes, one per channel. Those medians come from 13 LinkedIn captions, 19 Facebook captions and 15 Instagram captions.

Those numbers come with a caveat. The measurements were fit on summer captions I had approved plus eight of my own rewrites, 56 in all. The owner later sent some of those summer posts back as hard to follow. So the bands describe what my approved summer posts looked like. They can't say whether that shape was right. A measured profile tells you how a voice behaves. It never tells you whether the audience trusts it, and that's why the voice-fit score in the system I'm building advises and never blocks.

Three channel panels for one equipment dealer. The median sentence runs 12 words on LinkedIn, 10 on Facebook and 8 on Instagram, measured on 13, 19 and 15 captions. On all three, the median caption used a contraction everywhere one could go.

ONE EQUIPMENT DEALER, THREE CHANNELS

LINKEDIN 12

words in the median sentence

Contractions: used everywhere one could go

13 captions measured

FACEBOOK 10

words in the median sentence

Contractions: used everywhere one could go

19 captions measured

INSTAGRAM 8

words in the median sentence

Contractions: used everywhere one could go

15 captions measured

Fit on approved summer captions plus eight rewrites, 56 in all. The numbers describe how those posts behaved, not whether readers trusted them.

Same company, same products, three sentence shapes. An equipment dealer's approved captions ran 12 words a sentence on LinkedIn, 10 on Facebook and 8 on Instagram, and contracted wherever they could.

Some features of a voice are stable enough to count. When Frederick Mosteller and David Wallace worked out who wrote the disputed Federalist papers, they did it with the dullest words in the language. Their abstract notes that "filler words of the language such as an, of, and upon" provide "fairly stable rates" for a given writer, whatever the topic. Hamilton and Madison could argue about anything, and their rates for "upon" still gave them away. Contraction rate, function words and the spread of sentence lengths behave the same way in a client's writing, which is why they make good measurements.

Resemblance to an author is still a different question from whether a reader hears the post as coming from one of their own. Mosteller and Wallace could tell Hamilton from Madison. Their method couldn't say whether a shopkeeper in New York found either man convincing. That second question belongs to the audience.

One studio's copy shows both halves. My first round of AI-assisted copy for that studio sounded like AI, and the client couldn't say what was wrong with it. I measured the client's own copy across 241 pieces in nine families of metrics and rebuilt the copy against that [measured voice](/case-studies/voice-fingerprint-deliverable). One number carries the gap: my first round contracted 4% of its verbs, and the client's own copy ran between 59% and 94%. On a first measurement, the rebuild landed at 87%, inside the client's range. A complaint the client couldn't explain became a distance anyone could check.

Who's speaking is the input that went wrong for the dealer on launch day. My standard requires every claim to name who saw it, which is the right rule for an article like this one: when I tell you something, you should know where it came from. A writer following that rule literally, inside the dealer's own social posts, attached the owner's name and the interview date to sentence after sentence. A count of the posts lined up for launch that day found the pattern in 135 of 252.

The fix moved the source behind the post instead of deleting it. Here's the pair I wrote into the rule. Before: "[The owner] described the offer in one line in August: 'just show up, drop it on your job.'" After: "We show up and drop it on your job." Same fact, same source, still recorded where the team can check it. The company was now the one talking. The test I use for this is to ask whether a trade magazine could print the paragraph about the company without changing a word. If it could, the paragraph is in the magazine's voice, and the company is a bystander on its own page.

Before: a post reads "[The owner] described the offer in one line in August: just show up, drop it on your job." After: the post reads "We show up and drop it on your job," and the owner's name and the date sit in a source note below it.

ON THE CLIENT'S OWN PAGE, THE COMPANY SPEAKS

BEFORE

[The owner] described the offer in one line in August : 'just show up, drop it on your job.'

AFTER

We show up and drop it on your job.

SOURCE NOTE, FOR THE TEAM

Owner interview, August: 'just show up, drop it on your job.'

Kept on file. It doesn't run in the post.

On a client's own channel, the company is the speaker. The fact stays the same and the source stays on file for the team, moved out of the post and into a note behind it.

The rule runs the other direction in this article. Here I'm the byline, so naming a source in the prose is exactly right, and turning my evidence into an unattributed "we" would hide it. The same rule puts the source in opposite places on two channels, which is why the byline is an input, not a constant.

Building a profile for a new client starts smaller than it sounds. I read at least five pieces the client's audience already engages with, which tells me the vocabulary they expect, how long their sentences run and what kind of proof they trust. From that I write a short voice contract: who's speaking, to whom, about what, two words for the tone, one pattern this client must never use and one it must always use. That contract is the first version of the profile. Every approved draft and every rejected one adds to it, and the measured numbers come later, once there's enough approved work to measure.

With the five inputs set, the voice compiler runs its operations in order. It declares the target first: who's speaking, to whom, on which channel, and how formal, funny, personal and pointed the piece should be. It renders each part of the content map in that voice, for that channel, which is where a three-sentence LinkedIn paragraph becomes a two-line Instagram caption carrying the same point. It runs the shared minimums over the result, patterns first and a judge for what patterns can't decide. It scores the draft's voice fit against the reference set and reports the score without blocking on it. It counts what it kept, so a pass that flattened the client's quirks shows up. Then the client or someone from the audience reads it the way the audience would, and every change they make gets filed with its reason. When the same correction shows up twice, it becomes a rule in the client's file, and the next draft has to pass it.

Ten chiropractic clinics can share most of a page structure: who each service is for, what happens at a first visit, how to book. If two of their sites read alike, the voice compiler did nothing.

Ten chiropractors make the test concrete. Most of their content map is shared: who each service is for, what happens at a first visit, what the clinic can and can't promise, the booking path. The voice compiler is where they part. Say one clinic was founded by a former college athlete and most of its patients are weekend runners. Another sits in a retirement town and sees people whose backs have hurt for twenty years. Their service pages carry the same parts in the same order. They shouldn't share a single opening line, because the first clinic's patients describe their problem as a training setback and the second clinic's patients describe theirs as not being able to pick up a grandchild. If two of the ten sites read alike, the voice compiler did nothing.

## Every client gets the same minimums underneath

Under every client's voice sits a set of minimums I call writing floors. A floor is a requirement every sentence meets, whoever the client is and whatever the channel. My standard puts it as a dial and a wall: the voice is the dial, and the floors are the wall.

A few of the floors, to make that concrete:

- No em dashes, anywhere.
- No "it's not X, it's Y" framing, which argues with a belief the reader never held.
- No "the" in front of something the reader hasn't met yet, as in "the solution" on first mention.
- Every claim names who saw it: me, a named source, a study, a practice the reader can look up.
- Every headline is one complete thought a stranger could repeat after one read.

Writing the floors down turned out to be the easy part. Getting a model to obey them, even a model that had just written them, is where the process fell apart. The first full version of my own writing standard ran about 21,000 words, and it was drafted by a model that had spent a day defining what bad machine writing is. Its opening paragraph broke its own rule against the "it's not X, it's Y" shape. I caught that on first read, along with a fragment tacked onto the end of the same paragraph for rhythm.

So a separate pass ran the document's own rules over the document, one rule at a time, and logged every change. It made 72 edits to a file that defined every pattern it was fixing. Thirty of them were that same negative contrast. Six were sentences announcing their own significance, five were filler, and three each were figures of speech that signaled range without explaining anything, lines kept only for their rhythm, objects given agency, purple phrasing and internal vocabulary. The writer's own check had searched for two surface forms of the negative contrast and walked straight past the version that spans two sentences.

A model that spent a day defining bad machine writing drafted a standard against it, and a separate pass still made 72 edits to that standard. Thirty of them fixed the negative contrast the document banned.

That's the first lesson the floors taught, and it shapes every check I've run since. Knowing what bad writing is doesn't stop a model producing it. The method that removes it runs one rule per pass, by a reader who didn't write the draft, with a log of every change, because a pass that changes one thing can be checked and a rewrite that changes everything can't.

The second lesson came from testing lists of banned words and phrases, the tool I'd reached for first. An audit of my own case studies had produced 270 flagged lines, and 180 of them quoted the offending text cleanly enough to replay. When those 180 were replayed through the detectors I already had, the pattern lists fired on 15. Thirteen of those fifteen were a sourcing check that happened to trip, which wasn't the problem the auditor had named. The strongest detection engine I had fired on none of the 180.

A grid of 180 squares, one per flagged line replayed through existing detectors. Fifteen squares are lit for pattern-list hits, thirteen of them from an unrelated sourcing check. The strongest engine lit none. Every other square stays dark.

180 FLAGGED LINES, REPLAYED THROUGH MY DETECTORS

180 flagged lines replayed

15 fired a pattern list, 13 of them an unrelated sourcing check

0 fired the strongest engine

Empty squares: no detector fired.

An audit of my case studies flagged 270 lines, and 180 quoted the text cleanly enough to replay through my detectors. Pattern lists fired on 15, mostly for an unrelated reason, and the strongest engine fired on none.

The largest class of real problems in that audit, about 89 of the 270 lines, was a sentence pointing at itself: a line announcing that an insight is coming, a card naming its own category, a caption describing what the figure is instead of what it shows. No word list catches those, because the words are ordinary. The defect is in what the sentence does, so a vendor whose quality control is a banned-word list has automated the part of the job that catches the least.

So the rule list behind my checks has two engines. Patterns decide what a pattern can decide: an em dash, a banned transition, an uncontracted "do not" in a voice that contracts. A judge decides the rest, one narrow question at a time. Does this "we" presume a closeness the reader hasn't granted? Does this sentence say anything the paragraph didn't already say? Does this summary keep the body's condition?

That judge is a model, which might sound like it contradicts what I said earlier about models grading models. The difference is in the question. A model asked "does this sound human?" grades its own dialect and rates it highly. A model asked whether one specific sentence asserts what the reader is feeling has a narrower, checkable job, and each of those jobs gets tested before it's trusted. A judged rule only gets to block a draft after it reproduces known answers on examples it has never seen, the same discipline I describe in [evals and observability](/wiki/ai-and-technical-development/default/evals-observability). The statistical signs of machine prose, like even sentence lengths and the same few connectives on rotation, are measured and reported. No model is asked to judge them, for the reason the 88-versus-30 result made plain.

The judge also has to know what to leave alone. The keep signals from the fifteen-article review sit in the same rule list: a concrete scene, a candid setback, a fragment whose missing piece is obvious, an analogy that shows a mechanism. When a span matches one, the checker doesn't flag it. Without that, a checker trained to find defects finds them everywhere and sands the client's voice down to the model's average, which is the failure the whole system exists to prevent.

The last lesson was about where rules live. When I went looking, the banned-word list existed in at least five copies across my projects, in code and in prose, and the copies disagreed with each other. A correction made in one copy had no way to reach the others. I consolidated them into one list, running from mechanical tells to claims bigger than their evidence, plus the keep signals, limits on what a fix may change, and a short block per client recording what's licensed for that client's voice. Every check I'm building reads from that one list. A prose document can explain a rule; only the list defines it.

Three voice blocks sit on one shared base. A construction dealer speaks as the company, a runners' clinic hears patients describe a training setback, and a retirement-town clinic hears patients who can't pick up a grandchild. The base lists five floors every client shares.

THE DIAL, SET PER CLIENT

CONSTRUCTION DEALER

Speaks as the company. The owner's name is on file so the checker catches a borrowed byline.

We show up and drop it on your job.

RUNNERS' CLINIC, AN EXAMPLE

Patients describe their problem as a training setback.

RETIREMENT-TOWN CLINIC, AN EXAMPLE

Patients describe theirs as not being able to pick up a grandchild.

THE WALL, THE SAME FOR EVERY CLIENT

- No em dashes
- No 'it's not X, it's Y' framing
- No 'the' before something the reader hasn't met
- Every claim names who saw it
- Every headline is one complete thought

The voice is the dial and the floors are the wall. A dealer and two clinics set the dial three ways, and all three stand on the same five minimums, which no client's voice gets to move.

That per-client block is where the floors meet the voice compiler. One client's block notes that its pages speak as the company, with the owner's name listed so the checker can catch a borrowed byline. Another client's block licenses a few patterns the general floors forbid, because that client's accepted work uses them. The wall is the same for everyone, and each client gets a short, written list of exceptions built from their own approved pieces.

## Construction posts went through more than a dozen rounds

The dealer's posts are the strictest test I've run the method on, because the owner's rule is that a post his customers can tell was written by someone who doesn't understand their work is worse than no post at all. The readers are machine operators and contractors, many of them between 45 and 65 by the owner's account, often reading on a phone in the cab. When he went looking for a writer, his job post asked for a professional, industrial tone with no marketing fluff. My first note on the old copy asked for the same thing in my words: "We have to be able to speak direct like an actual construction operator out on a job site, or if some gentlemen ran into each other at a bar, how would they speak?" The platform side of that engagement, where an agent platform drafts posts from vendor brochures and spec sheets, recorded interviews and support stories, is written up as a [case study](/case-studies/contentfactory-positioning-dealership). The social rounds are the part it doesn't cover.

The owner's first feedback came on a batch of summer posts. He approved five and sent twelve back with notes. Six of the twelve said he didn't understand a sentence. Five of those six landed on the third sentence or the third paragraph. My reading of that pattern is that sentence two wasn't carrying the reader into sentence three, so by the third line he was guessing. The rule that came out of it: every sentence has to carry the one before it forward, and confusion that starts small has compounded by line three.

A timeline of rounds on a dealer's social posts: summer send-backs, a launch-day byline sweep of 135 of 252 posts, 156 calls to action rewritten, new post rules, majors cut from 10 of 33 to none, write passes shrinking from 33 captions to 1, and a queue waiting on the owner.

MORE THAN A DOZEN ROUNDS, ONE CHANGE EACH

- SUMMER SEND-BACKS 5 approved, 12 sent back. 6 of the 12 said a sentence didn't make sense.

- LAUNCH DAY A borrowed byline in 135 of 252 queued posts.

- CALLS TO ACTION 156 asks rewritten.

- POST RULES A subject and a verb in every sentence, the four-line brief, nine named slop patterns.

- REVIEWER ROUND 10 of 33 captions rated major. None after one revision. 10 0

- WRITE PASSES Fewer posts needed rewriting each pass. 33 14 7 2 1 CAPTIONS PER PASS

- A LATER BATCH Majors fell from 8 to 3 after its rewrite. 8 3

- THE QUEUE Calendar pushed back twice. The posts wait on the owner's read.

Each round on an equipment dealer's posts changed one thing and kept the change as a rule. Major problems fell from 10 of 33 captions to none, and each write pass touched fewer posts than the last.

Launch day added the borrowed-byline rule from the previous section, after 135 of 252 queued posts turned out to quote the owner like a news source, or cite the interview, on his own page.

The week after launch was the densest. It opened with a round on the calls to action. The drafts kept ending with the shop's phone number, and my note on that was blunt: putting the number in "is insulting and it is silly. Most of these people already know" the owner. The asks got rewritten to things an operator would act on, like swinging by the shop for a demo. That round rewrote 156 of them.

Then came the post rules, because clean sentences were still producing posts I called slop. The first rule is mechanical: every sentence in a post has a subject and a finite verb. "20 minutes. Every crew. Every morning." fails, because a cold reader can't tell what the twenty minutes is doing. The second rule is the four-line brief from the content section, with the point stated first and stated as a benefit the viewer gets.

The third rule became a list. Every time I said "that's slop" about a post with clean grammar, the pattern behind the verdict got found, named and written down, so every later post could be checked against it. Nine patterns came out of that week. A line like "Five crews lose 166+ paid hours a season" presumes the reader runs five crews; he might run one. A post opening "One tiltrotator can work on more than one of your excavators" answers a question nobody asked yet. A line like "When crews decide, they want it yesterday" narrates the seller's private knowledge about how buyers buy, to the buyers.

Nine cards, one per named pattern, each with an example from a real draft: presumed reader, answer with no question, purpose stated last, transcript stitching, caveat that talks down the product, disconnected assertions, no stated benefit, throat-clearing lead-in and leaked sales playbook.

NINE NAMED PATTERNS FROM ONE WEEK OF POSTS

01 Presumed reader

"Five crews lose 166+ paid hours a season."

PASSES: Frame it as a hypothetical: "Say you run five crews."

02 Answer with no question

"One tiltrotator can work on more than one of your excavators."

PASSES: Open on the buyer's real question.

03 Purpose stated last

"click one button, pull back on the stick, machine control cuts grade"

PASSES: Say what it's for, then how.

04 Transcript stitching

"is for light material placement like that"

PASSES: Keep the client's facts and write new sentences.

05 Caveat that talks down the product

"a huge highway job still needs a motor grader"

PASSES: Say what the product is for.

06 Disconnected assertions

A list of attachments with no reason to care.

PASSES: Tie the list to the reader's use.

07 No stated benefit

"I don't understand the benefit of this."

PASSES: State the benefit before the ask.

08 Throat-clearing lead-in

"With machine control, the machine…"

PASSES: Put the subject and verb first.

09 Leaked sales playbook

"When crews decide, they want it yesterday."

PASSES: Say what the product does and what the viewer gets.

Every post I called slop had clean grammar, so each verdict got traced to a pattern, named and written down. Nine came out of one week, and every later post was checked against all nine.

The remaining drafts then went through the rounds that produced the numbers in the opening. A reviewer model that hadn't written any of them graded the first full rewrite of 33 captions and rated 10 as major problems. After one revision, it rated none that way. Four of those ten had brought back wording I'd already cut in an earlier round, which is the repeat-correction problem in its plainest form, and fourteen of the 33 calls to action ended on the same word, "yourself."

The reviewer checked facts as well as wording. On a later batch, it caught a post that described a trench job as one person's work from the cab, and the owner's own recorded words said otherwise: one of the two people in the trench is freed up, and the other stays to push the joint together. It also caught three near-copies of posts already scheduled, down to whole sentences.

Each of my edits became a named rule, and every remaining post had to pass it. A few of them, in the order they showed up:

One end card told operators to put their own hands on the sticks. Read by a heckler, it says something else, and I couldn't publish it. The fix invited them to book a demo and run a cut themselves. The rule became a heckler test, and a sweep of 297 captions found that line was the only instance.

The remote-support post from the content section was the next. My reaction to its opener: "I have to see myself as a helpless victim that's requiring someone to drive out before I even know any context on it." The rule: open on the situation, and let the trouble belong to the machine, not to the person reading.

Each edit I made to an equipment dealer's posts became a named rule that every remaining post had to pass. One sweep checked 297 captions against a single new rule and found one line to fix.

A post corrected a belief the reader might hold, and it came out confrontational. I compared it to "trying to convince someone to read the Bible by beating them with it." The rule: offer relief, and never tell the reader what he believes.

A post listed a feature and stopped. My note: "The 'so what' has to be the central focus of ANY post." Another said the reader would have someone to call, and the rule that came out of it was a question every post now gets asked: what, specifically?

Some of my edits were single words. "Model" became "specifications," because that's the buyer's word for it. A support line became "supports your metal remotely," because "metal" is what operators call their machines.

Others were about the whole post. A draft built around a statistic on operator job openings, most of them replacements, stated the number and stopped there. My note was a sequence: "What is the point you're trying to make? ... Start from there, then write it, then ask yourself if what you wrote meets the point." A Facebook post came back from me marked "written like a tweet," the channel profile from the voice section failing in a single post. And one post lost its meaning in a shortening pass, because the cut removed the one step that connected the problem to the product.

One post shows most of this at once. Its first draft opened, "Every crew spends 20 minutes setting up a base every morning." That sentence states a fact about every reader's crew. The reader might not run a base station at all, and the line reads as the dealer telling him about his own morning. The rewrite asks: "Does your crew spend 20 minutes every morning setting up a base?" Then it states what the product does: it skips that setup by sending corrections over a cell connection, so there's no base on a tripod. Then the ask.

In August, the owner had approved an earlier version of the same idea that opened "20 minutes setting up your base. Every morning. Every crew." On that line, my floors are stricter than my client. His approved post breaks two of my rules, the presumed reader and the verbless sentence, and he approved it. I kept the rules anyway, because the rewrite carries the idea and the numbers he approved without the presumption. It's also the reason his ear stays in the loop: the floors can be wrong about what his audience accepts, and only his read or theirs settles it.

A four-line brief. Audience: crews that set up a base station every morning. Point: you can skip that setup. Proof: corrections arrive over a cell connection, so there's no base on a tripod. Ask: book a demo. Below it sits the post's opening question.

THE FOUR-LINE BRIEF, FILLED IN

AUDIENCE Crews that set up a base station every morning.

POINT You can skip that setup.

PROOF Corrections arrive over a cell connection, so there's no base on a tripod.

ASK Book a demo and see a receiver pick up corrections with no base on a tripod.

THE OPENING LINE IT PRODUCED

"Does your crew spend 20 minutes every morning setting up a base?"

Before a dealer's post about base setup got rewritten, it got a four-line brief: who it's for, what they get, why they should believe it, and what to do next. The opening question came out of that brief.

That's where the work stands. The rounds took the drafts through every check a rule or a cold reader can run. The calendar was pushed back twice so the owner would have time to read the queue properly, and the posts wait on him. He told me before he read them that he'd skip the grammar and read for whether they make sense and sound right. The rounds have done what checks can do.

## Six kinds of buyers hit the same split in different places

Partway through my work for the AI-tools publisher, I put one line at the top of the file where I keep that client's corrections: "If I have to go and make the same correction every single time, then I might as well just write it myself." That line is the economics of every engagement below. The file of corrections, each with its reason, is what stops the second fix, because every correction becomes a rule the next draft has to pass.

The situations below come from job posts I've read, and where a buyer's own words say it best, I've quoted the post.

### An agency writes for a whole niche of clients

A shared page map with five parts sits across the top: who it's for, the problem in the patient's words, what the clinic does, the proof and the ask. Below it, three clinics each keep a voice profile and a file. One owner's edit goes into that clinic's file only.

SHARED PAGE MAP, BUILT ONCE FOR THE NICHE

- Who it's for
- The problem in the patient's words
- What the clinic does about it
- The proof
- The ask

CLINIC A

Voice profile from its reviews, its front desk, the founder's recorded talk and its patients' questions

THIS CLINIC'S FILE

CLINIC B

Voice profile from its reviews, its front desk, the founder's recorded talk and its patients' questions

Owner edits a draft

THIS CLINIC'S FILE Stays out of the shared rules

CLINIC C

Voice profile from its reviews, its front desk, the founder's recorded talk and its patients' questions

THIS CLINIC'S FILE

An agency writing for a niche builds the page map once and reuses it for every clinic. Each clinic's voice and edits stay in that clinic's file, because one owner's taste isn't a law for the other nine.

One post asked for a writer to "produce full landing page copy for multiple telehealth clinics." Another came from a company writing for HVAC owners, plumbers and roofers. That's the ten-chiropractors problem with real job titles on it, and it's the case the split between content and voice was built for.

The content compiler builds one map per page type and reuses it across every client in the niche. A service page gets the same parts each time: who the service is for, the problem in the patient's own words, what the clinic does about it, the proof, the ask. Each part carries its job, so a writer filling in clinic number seven knows what the proof block has to accomplish before writing a word of it. That shared map is the agency's process, and it's the part that should never be rebuilt per client.

The voice compiler is where the clinics part ways. Each client gets a profile built from that client's own material: its reviews, the way its staff answer the phone, the founder's recorded talk, the questions its patients ask. Each channel the clinic publishes on gets its own profile on top. The check is simple: put two client pages side by side, read the openings aloud, and listen for which clinic is talking.

One rule keeps a ten-client system from drifting. When a clinic owner edits a draft, that edit goes into that clinic's file. It doesn't go into the agency's general rules, because one owner's taste isn't a law for the other nine. I learned that with the dealer, whose notes are his company's standard and stay there.

### A founder's posts go out under the founder's name

A founder's calls supply the facts, opinions and stories worth posting. The spoken wording stays behind, because a phrase pasted from a transcript reads as stitched, and the sentences get written fresh in his measured voice.

The ghostwriting posts were specific about what they feared. One asked for "posts in the founder's voice, not yours," built from founder interviews and sales calls. Another wanted three LinkedIn posts and two X threads a week in the founder's voice, and offered a voice pack with his published posts, approved scripts and rejected drafts, each with the reason it was killed. That post said most of the job would be editing and rejecting.

Here the content compiler decides what's worth saying this week, and the founder's calls are the best source for it. The calls supply facts, opinions and stories. They don't supply wording. A spoken phrase pasted into a post because it's sourced reads as stitched: I've seen a post carry "is for light material placement like that," where the "that" pointed three sentences back into a conversation the reader never heard. Facts come from the transcript, and the sentences get written fresh.

The voice compiler needs a profile of how the founder sounds when nobody's editing him. Mine was built from about 115,000 words of my own transcribed speech, and it measures things like how often I contract, how long my sentences run before I end them, and which words I never use. A founder's profile gets built the same way, from unscripted calls, the way I measured the studio's own copy before rebuilding mine. The rejected drafts with their reasons can teach more than the approved ones, because each reason is a rule waiting to be written.

The founder's ear closes the loop. Every post he kills goes into his file with the reason he gave, and when the same reason shows up twice it becomes a rule.

### A brand has a voice guide that hires keep getting wrong

Four lines from a trades branding firm's voice guide, each tagged by who enforces it. No em dashes goes to the checker. Customer framing with no owner blaming goes to the judge. Operator tone goes to the judge and the editor's ear. No one-word sentences goes to the checker.

ONE BRANDING FIRM'S VOICE GUIDE, LINE BY LINE

CHECKER a pattern decides JUDGE a narrow model call, trained on the firm's approved examples EAR only the firm's editor can settle it

"No em dashes" CHECKER

Decided every time.

"Customer-framed, never owner-blaming" JUDGE

Needs approved examples of each side before it gets a vote.

"Operator tone" JUDGE EAR

A model flags corporate filler. The editor decides whether it sounds like someone who's run a crew.

"No one-word sentences for punch" CHECKER

Overrides my default, which keeps a deliberate fragment. Recorded in the firm's block.

A voice guide becomes a specification once each line has an enforcer. A pattern decides some lines, a narrowly trained judge decides others, and the firm's editor keeps the call on whether a sentence sounds like someone who's run a crew.

The post closest to the dealer's world came from a branding firm that writes for trades and construction contractors: "Our content lives or dies on voice, and most writers get it wrong on the first try." The firm had already written its voice down as rules: no em dashes; "customer-framed, never owner-blaming"; "operator tone," with no corporate fluff or motivational filler; "teaching, not selling"; and flowing prose with no one-word sentences for punch.

For a buyer like that, the content compiler changes little. The firm already knows what each piece is for: every piece teaches one idea an owner can use.

The voice compiler's job is to turn each line of the guide into something checkable, and the lines fall into three kinds. "No em dashes" is a pattern; a checker decides it every time. "Customer-framed, never owner-blaming" is a judged rule, and the judge needs examples of each side from the firm's own approved pieces before it gets a vote. "Operator tone" is partly judged and partly ear: a model can flag corporate filler, and only the firm's editor can say whether a sentence sounds like someone who's run a crew.

One line in that guide overrides a default of mine. My floors protect a deliberate fragment as part of a writer's voice. This firm bans them outright. The firm wins on its own pages, and its block in the rule list records that exception, so the checker flags a fragment there that it would leave alone anywhere else. When the editor marks a sentence that drifted, the mark goes into the firm's file with the reason, and the guide gets a little more checkable every round.

### An operator already wired AI into the brand

One buyer's post asked for someone to decide what stays, what changes, what needs refinement and what never ships. That's a verdict per part, and every draft pulled from the batch carries its reason.

Some buyers are past the question of whether to use AI. One post described a company that had already loaded its brand voice, customer avatars, testimonials and podcast content into its tools, and it was hiring someone to decide "what stays, what changes, what needs refinement, and what should never make it out the door." The same post said it didn't want someone who copies and pastes AI output.

That sentence is the content compiler's verdict step, written by a buyer in his own words. Stays, changes, needs refinement and never ships map onto keep, fix, rework and cut, made one part at a time and recorded with a reason. A batch run through that step comes back with its verdicts visible, so the next batch can be briefed against them.

The voice side shows up in a different post from the same kind of buyer, one hiring for ad scripts narrated by several different people. It asked how a writer keeps each narrator's voice distinct when the drafting tool falls back on the same rhythms and phrasing every time. That's a voice compiler with several profiles running under one brand. Each narrator gets a profile of his own: the words he'd use, the length of his sentences, the proof he'd reach for, the jokes he wouldn't make. The drafts get scored against each narrator's reference set, so a script that drifts toward the tool's default rhythm shows up as a distance from that narrator, before anyone records it.

The architecture behind this is the per-part checking from the original content compiler entry: when one part fails, only that part goes back for rework, and the rest of the batch stays put.

The person signing off on the batch has the job the system can't do without. Whatever that person approves becomes the reference set, and whatever they reject, with the reason, becomes the next rule. A buyer in this position already has the generation side built. What they're missing is the record of decisions that turns a reviewer's taste into something the next batch can be checked against.

### One brand talks to several audiences on several channels

Two messages from an invented coffee roaster under one band that reads same voice: dry, plain, no exclamation marks. A restock text says the Ethiopia is back from the same farm. A win-back email mentions a changed blend and offers the next bag free.

SAME VOICE Dry, plain, no exclamation marks.

INVENTED EXAMPLE

RESTOCK, TEXT MESSAGE

The Ethiopia's back. Same farm, same roast. It went fast last time.

JOB: Say it's back and get out of the way.

WIN-BACK, EMAIL

It's been a while since your last bag. We changed the house blend this spring and we'd like to know what you think, so the next one's on us.

JOB: Give a lapsed customer a new reason to come back.

An invented coffee roaster with a dry, plain voice sends two messages. The restock text only says the coffee's back, and the win-back email carries a new reason to return.

Lifecycle buyers run flows: a welcome series, a restock notice, a win-back, a review request, each going out by email, text message or push notification to a different group of customers. One retention post put the whole problem in a phrase: "a restock notice reads differently than a win-back." The brand's voice was already documented. The job was making every send sound like that brand while each one did a different job.

The content compiler asks whether each message earns its place in the journey and what its one job is. A win-back has to give a lapsed customer a reason to come back that wasn't there when they left. A restock notice has to tell someone who wanted a product that it's back, and get out of the way. A message that can't state its job in a sentence is the first cut.

The voice compiler holds the brand voice fixed and moves the tone per channel and per stage, which is that style guide's line about voice and tone doing practical work. A text message gets the brand at its most conversational. A win-back email gets the brand admitting it noticed you left.

Here's an invented pair to show it, for a coffee roaster whose voice is dry and plain. The restock text: "The Ethiopia's back. Same farm, same roast. It went fast last time." The win-back email: "It's been a while since your last bag. We changed the house blend this spring and we'd like to know what you think, so the next one's on us." The two share one dry, plain voice with no exclamation marks. The restock only has to say the coffee's back. The win-back carries the new reason to return: a changed blend and a free bag.

The customers' own words do a lot of this work. In an older engagement for a supplements brand, the email work mined customer reviews and competitor feedback to write [in the language buyers used](/case-studies/fitness-supplements-email-ppc), and a sequence with no hard pitch ran before the one that asked for the sale.

### A content engine has to take positions

Editorial buyers want argument. One post for an interview-based writing role listed what it wanted in plain terms: "Taking a position and defending it. We don't want both-sides hedging." The same post asked for every statistic to be traced to its original source, dated. Several posts in this group also ban generic AI language outright.

A strong claim backed by a named source can stay strong. Stack 'may,' 'might' and 'can' under an overreach and it still wobbles; name the condition under which it holds and the claim keeps its edge.

Those two demands look like they pull against each other, a strong position and a careful source, and the content compiler is where they stop pulling. The claim checks from the fifteen-article review do most of the work. A capability doesn't get stated as a guarantee. An anecdote doesn't get stretched into a law. A summary keeps the body's condition. Every claim names who saw it. None of those checks forces a hedge. The review's own advice was to keep strong claims "when they are warranted," and to replace an overreach with the condition under which the claim holds, not with "may" and "might." A position backed by a named source can stay sharp, and that's the version an editorial buyer is paying for. The grounding line from the content section is the pattern: the overreaching version said grounding makes an answer checkable, and the repair kept the claim and added the condition that makes it true, "when the system exposes its sources."

The voice compiler gives each channel its own editorial voice. A newsletter can carry a longer argument and a dry aside. A LinkedIn post from the same publication makes the claim in its first sentence and the case in three more.

The editor's ear produces some of the most useful data in this setup, and the fifteen-article review explained why. Rejected suggestions belong in the record with their reasons. If the checker flags a joke because it read it too literally and the editor keeps the joke, that rejected flag teaches the system something accepted edits can't: where its judgment is wrong.

## Running the method by hand proved some things and left a gap

I've said "by hand" throughout. There's no such thing as a content compiler yet, as software I could hand you. What exists is the method, run on real client work with models drafting and every operation checked, to find out whether the quality is reachable at all before anyone builds the software that runs it.

Three parts are proven. The first is that the shared minimums can be reached on real volume: on one set of 33 of the dealer's captions, the reviewer's major flags went from 10 to none in one revision, and a later batch went from 8 majors to 3 after its rewrite. The second is that per-client rules accumulate and stop repeat corrections. The file of standing rules for the AI-tools publisher holds 43 of them, built up over eight assignments, and each one is a correction I no longer make by hand. The dealer's nine slop patterns and the rules from my own edits did the same job across his queue, and the write passes on his remaining drafts shrank from 33 captions to 14, then 7, 2 and 1.

Three columns. Proven: major flags cut from 10 of 33 captions to none, 43 standing rules for one publisher, and drafting first with at most three questions. Not yet proven: whether a measured voice profile carries the audience's voice. Still to build: the checker, readable profiles, edit capture and the software.

PROVEN

Shared minimums on real volume. Major flags went from 10 of 33 captions to none in one revision, and a later batch from 8 to 3 .

Per-client rules stop repeat corrections. One publisher's file holds 43 standing rules from eight assignments.

Draft first, ask second. A marked gap for each missing fact, and at most three questions.

NOT YET PROVEN

Whether a measured voice profile carries the voice an audience hears as one of its own.

STILL TO BUILD

The checker that reads the rule list. The 55 rules are written down, and no automated check reads them yet.

Voice profiles a system can read. Today they're written rules, kept by hand.

Capture of every edit. Today edits get captured and turned into rules by hand.

The content compiler as software a client can log into.

Running the method by hand proved three things and left one question open. The rule list, the voice profiles and the edit record exist on paper and by hand, and the software that runs them is still being built.

The third is drafting first and asking second. When a draft needs something only the client knows, like a story from his own jobs or a number from his books, the draft leaves a marked gap where that material goes, and the client gets at most three questions, each tied to a specific gap. A client's attention is the input I can least afford to waste. A marked gap costs one answer to fill. An invented detail costs his trust when he finds it.

Some of it isn't proven. I don't yet know whether a measured voice profile can carry the voice an audience hears as one of their own. The floors measure what's wrong with a sentence. The voice-fit score measures how close a draft sits to the client's approved pieces. Neither one measures whether an operator reading on his phone hears someone who's run a machine. Bell's research says why that's a separate question: style is a response to an audience, so the audience is the only one who can say whether the response landed. Mosteller and Wallace showed what can be counted. Counting can't say whether a community trusts the voice.

So the method treats the audience's ear as data to collect, not a rule to write in advance. Every reaction to a real draft gets recorded: what sounded off, in the reader's words, and why. Approved pieces go into the reference set with a note on what makes them work. The voice profile is designed to be re-derived from that evidence every few dozen records, so it follows what the audience accepts instead of what I guessed they would.

Four steps in a row. A client edits a sentence and gives a reason. The edit waits in that client's file. A second edit arrives with the same reason. It becomes a rule, and the next draft is checked against it and passes.

HOW A REPEATED CORRECTION BECOMES A RULE

EDIT AND REASON

A client changes one sentence and says why.

CLIENT FILE

The edit waits in that client's file.

SAME REASON AGAIN

A second edit arrives with the same reason.

NEW RULE

The next draft gets checked against it and passes.

Every edit goes into the client's file with its reason. When the same reason shows up a second time, it becomes a rule the next draft has to pass, so nobody makes that correction by hand again.

The ear doesn't have to belong to the client. For the dealer, the ideal reader is an operator who's never seen the brief, reading on his phone in the cab. The owner is the closest reader available, and by the rule I set for his notes, they carry how operators talk and what lands in the trade, not grammar or timing. For a clinic, the ear is a patient's. For a founder, it's the founder's own, because his audience is reading him. Whoever holds the ear, their edits go into the file with a reason, and the profile follows them.

There's also a gap between the method and a system, and I'd rather a buyer hear it from me. The rule list exists: 55 rules, designed and written down, with the keep signals and the per-client blocks. No automated check reads it yet, and the checker that will is being built. Voice profiles exist as written rules per client, kept by hand. Edits get captured by hand and distilled into rules by hand. The content compiler that judges each part's value is a set of operations I run, with models doing the drafting and a separate reviewer doing the grading, and it isn't yet software a client can log into.

The way it scales once it's software is already designed. The writing method lives in one kit per role: the operations, the reference material, the file of corrections and the instructions for each tool. That kit is the same for every client. What changes per client is the voice profile, the reference set and the client's own file of edits, and those get swapped in. Ten chiropractors would share one copywriting kit and carry ten profiles. It's the compiler picture from the original entry again, one toolchain with many targets, applied to voices instead of machines.

What a buyer gets today is the method run on their own work, round by round, with every correction kept and turned into a rule, and the software arriving underneath it as each piece gets built. If you need a tool you can install this week and run without anyone in the loop, this isn't that, and I'll say so on the first call.

## Four questions sort a system from a prompt

Whoever writes for you, whether that's me, an agency or a tool, four questions show whether there's a system behind the copy or a prompt and a hopeful reviewer.

A card of four questions for a writing vendor, each with a good and a weak answer: the parts and their jobs, where your voice lives in writing, the minimums every sentence meets, and what happens to your edits.

FOUR QUESTIONS FOR WHOEVER WRITES FOR YOU

ASK A SYSTEM SHOWS YOU A WEAK ANSWER

What are the parts of this piece, and what's each one's job?

A SYSTEM SHOWS YOU

A content map: the parts, their jobs, and which got cut and why.

A WEAK ANSWER

One long draft with no parts named.

Where does my voice live in writing, and can I read it?

A SYSTEM SHOWS YOU

A written voice profile you can open.

A WEAK ANSWER

'In the prompt,' or 'in our writers' heads.'

What minimums does every sentence meet, and how are they checked?

A SYSTEM SHOWS YOU

Named checks, including what catches a claim stronger than its evidence.

A WEAK ANSWER

A banned-word list and nothing else.

What happens to my edits?

A SYSTEM SHOWS YOU

A file of your edits with reasons, turned into rules.

A WEAK ANSWER

Corrections that vanish into the next draft.

Ask these of any writer, agency or tool. A system shows you a content map, a written voice profile, named checks and a file of your edits, and a prompt with a hopeful reviewer can't.

- What are the parts of this piece, and what's each one's job? A vendor with a content compiler, by any name, can show you the map: the parts, their jobs, and which ones got cut and why.
- Where does my voice live in writing, and can I read it? If they say "in the prompt" or "in our writers' heads," the voice will drift the first time the writer or the model changes.
- What minimums does every sentence meet, and how are they checked? A list of banned words is a start. Ask what catches a claim stronger than its evidence, or a summary that drops the body's condition.
- What happens to my edits? If your corrections disappear into the next draft with no record, you'll be making them again next month.

If you publish five pieces a month on one channel and you're happy editing them yourself, a good writer with a good prompt is the right tool, and my original content compiler entry says the same: below a certain volume, the architecture costs more than it saves. If you need a turnkey tool with nobody in the loop, I've told you above where the line sits today. And if nobody has ever sent back one of your drafts saying it didn't sound like them, you probably don't have the problem this solves.

The buyers this is for have a different week. They're publishing for many clients, or many audiences, or under someone else's name, and they're fixing the same sentence for the third time. For them, the split pays off in the one place they can see it: the corrections stop repeating.

The bolt of amber cloth and three dress forms return with labels. The long overcoat is LinkedIn at 12 words a sentence, the jacket is Facebook at 10, the vest is Instagram at 8, and the stone wall behind them is the shared minimums.

LINKEDIN the long cut: 12 words a sentence

FACEBOOK 10 words a sentence

INSTAGRAM the short cut: 8 words a sentence

Median sentence lengths from one equipment dealer's approved captions.

THE CLOTH one content map

EACH CUT one client, one audience, one channel

THE WALL shared minimums no voice moves

THE EAR a reader from the audience decides whether it sounds right

The same cloth is one content map. Each cut fits one client, one audience and one channel, the wall of minimums stays put behind them, and a reader from the audience decides whether it sounds right.

Go back to the ten chiropractors. They share one map: who each service is for, what happens at a first visit, what a clinic can promise, how to book. Above that map sit ten voices, one per clinic, each built from that clinic's own patients and its own words. Underneath sits one set of minimums that none of them gets to break. And at the end of every round there's a person like the owner who told me he wouldn't read for grammar. He'd read for whether it sounds right.

Source: https://andydataguy.com/wiki/ai-and-technical-development/default/voice-compiler

