Most AI writing advice begins with a blank chat window and a carefully engineered prompt. That starting point skips the part of the process where the actual thinking usually happens. I begin with a voice memo. Not because it is fashionable, but because speaking first forces a different kind of clarity than typing.
This is the full workflow I use to take a rough spoken idea all the way to a finished, publishable article. It is not a theoretical framework. It is the sequence I follow when the work has to hold up under real editorial standards. AI is present at several stages. Human direction remains responsible for structure, voice, and the final decision about what is worth publishing.
Why Start With a Voice Memo
Writing directly into a prompt often produces language that sounds finished before the thinking is finished. The model fills gaps with fluent defaults. The result can look complete while still lacking a clear point of view.
A voice memo changes the sequence. I walk, or sit with the question, and speak the idea in the order it actually arrives. The recording is usually messy. Sentences restart. Digressions appear. The central claim is sometimes not obvious until the end of the memo. That messiness is useful. It preserves the shape of the original thinking before it is smoothed into professional language.
I keep the memos short. Ten to fifteen minutes is usually enough. Longer recordings tend to circle. The goal is not a complete draft in spoken form. The goal is a raw record of the argument as it exists before any tool begins shaping it.
Stage 1: Capture and Transcribe Without Cleaning
I record the memo on my phone and transcribe it with minimal intervention. I do not ask AI to clean up the language at this stage. Cleaning too early removes the hesitations and repetitions that often signal where the real emphasis belongs.
The raw transcript becomes the source material. I read it once through and mark the moments that feel charged or precise. These marked passages become the candidates for the core of the article. Everything else is treated as supporting or disposable.
This stage is deliberately low-tech. The only decision that matters is which fragments already carry weight.
Stage 2: Extract Structure Before Generating Prose
Once the strong fragments are identified, I use AI to help propose structure. I do not ask for a full draft. I ask for possible outlines that organize the marked ideas into a coherent sequence. I usually request three different architectural options: one chronological, one argument-led, and one that begins with the strongest claim.
I review the outlines against a simple standard. Does this structure make the central point unavoidable? Does it create unnecessary delay before the reader understands what the piece is doing? Does it force repetition?
Most of the generated outlines are discarded. The surviving structure is rewritten in my own words until it feels accurate. Only then do I move to drafting.
This separation is important. Generating structure and generating prose are different tasks. Mixing them usually produces writing that is fluent and loosely organized rather than deliberate.
Stage 3: Draft in Controlled Passes
I draft in short, controlled passes rather than asking for a complete article in one step.
First pass: expand the outline into rough sections using the language from the original voice memo wherever possible. The goal is to preserve the spoken cadence and the original emphasis.
Second pass: use AI to generate alternative versions of weak sections. I feed it the surrounding context and the specific problem with the current version. I do not accept the first alternative. I generate several and then choose or combine.
Third pass: human rewrite for voice and precision. This is the longest stage. I cut hedging language, replace generic transitions, and restore any specificity that the model has smoothed away. I also check every claim against the original memo to make sure nothing essential has been lost or distorted.
The drafting process is iterative and intentionally incomplete until the final pass. Treating any AI output as a finished section is the fastest way to end up with competent but impersonal writing.

Stage 4: Edit for Decision, Not Just Polish
Editing is where most of the value appears. I edit in layers.
The first layer is structural. Does every section earn its place? Can any section be cut without weakening the argument? Are the transitions doing real work or simply filling space?
The second layer is voice. I read the draft aloud and mark every sentence that sounds like it could have been written by a model trained on average professional content. Those sentences are rewritten or deleted. The goal is not to make the writing sound casual. The goal is to make it sound like a specific person who has a point of view.
The third layer is precision. Vague intensifiers, empty transitions, and claims that sound stronger than the evidence is removed. I also check for places where the writing has become more abstract than the original voice memo. Abstraction is often a sign that the model has filled a gap rather than that the thinking has become clearer.
Only after these layers do I consider surface polish: rhythm, sentence variety, and final word choice. Polish applied too early hides structural problems.
What I Consistently Cut
Certain patterns appear in almost every AI-assisted draft. I cut them on sight.
Opening paragraphs that restate the topic without making a claim.
Transitions that announce structure instead of advancing the argument.
Sentences that hedge with “in today’s world,” “it is important to note,” or similar empty framing.
Conclusions that summarize what has already been said instead of leaving the reader with a sharper final thought.
Any passage that sounds generally true but could apply to many other articles on the same subject.
Cutting these patterns is not a matter of style preference. It is a way of protecting the specificity that the voice memo originally contained.
Building a Repeatable Rhythm
The workflow only becomes efficient when it is repeated. I keep a short personal checklist that travels with every article:
Has the original spoken emphasis survived?
Is the structure serving the claim or merely organizing information?
Where has the language become generic, and can those sentences be cut or rewritten?
Does the finished piece still feel like it came from the same person who recorded the memo?
Over time the checklist becomes internalized. The early stages move faster because the decision criteria are already clear. The later stages remain slow because judgment does not compress as easily as generation.

How This Workflow Differs From Prompt-First Writing
Prompt-first writing treats the model as the origin of the text. The human role becomes one of refinement and selection. The risk is that the final piece inherits the model’s average sense of what an article on the topic should sound like.
Voice-memo-first writing treats the model as an accelerator and a stress-tester. The origin of the thinking remains human. The model helps organize, expand, and challenge that thinking. The final responsibility for what is said and how it is said stays with the writer.
The difference shows up most clearly in the finished voice. Prompt-first pieces often feel competent and slightly anonymous. Voice-memo-first pieces are more likely to retain the particular rhythm and emphasis of the person who made the original recording.
Practical Constraints and Limits
This workflow assumes that the writer already has something to say. AI cannot supply the underlying observation or judgment. It can only help develop and pressure-test what is already present in the voice memo.
It also assumes a willingness to delete a large percentage of what the model produces. Writers who treat every generated paragraph as potentially usable will end up with longer, softer drafts. The discipline of rejection is part of the process.
Finally, the workflow is slower than asking for a complete draft in one prompt. The extra time is spent on decisions that determine whether the article is worth publishing. In practice the overall timeline is often shorter because fewer major revisions are required after the first full draft exists.
Closing the Loop
When the article is finished, I sometimes return to the original voice memo and listen to it once more. The comparison is instructive. The finished piece should feel like a clearer, more deliberate version of the same thinking, not like a different piece of writing that happens to cover the same topic.
If the finished article has lost the charge that was present in the spoken version, the editing process has gone too far in the direction of conventional polish. If the finished article is still as diffuse as the original memo, the structural and decision stages were not rigorous enough.
The voice memo is the beginning. The finished article is the result of a series of deliberate choices about what to keep, what to cut, and what to insist upon. AI can accelerate several of those stages. It cannot make the choices. That part remains the work.
Travellers Write
No letters yet — be the first traveller to write.