Back to Blog

Eleven Months of Talking to AI Agents: What 2.26 Million Dictated Words Show

I do not type much. I dictate almost everything I give an AI agent, using a tool called WhisperTyping that transcribes as I speak and keeps a copy of every recording. I had never really looked at that archive. Last week I went digging through it for something else entirely and realised I was sitting on a fairly unusual record: eleven months of one person talking to AI coding agents, nearly every day, timestamped.

So this is what came out of it. 14,625 recordings, 2,258,517 words, 332.7 hours of audio, from 26 August 2025 to 25 July 2026, across 292 active days out of 334. All of it me dictating prompts while building EF-Map and a few other EVE Frontier tools. Some of what I found matched what I would have guessed. A few things were the opposite of what I would have sworn was true.

How to read all of this

Every count is a literal string match over raw transcripts, so paraphrases are missed and transcription errors are included. Monthly charts run September 2025 to July 2026: August 2025 is a six-day stub, and July 2026 covers 1 to 25 July, which is why the volume chart is per active day rather than per month. The hour-of-day and ask-position charts use all 14,625 recordings. And in the spirit of the subject, I did not write the analysis code or the charts by hand. I dictated what I wanted at Claude and it wrote the Python and the SVG. Each chart has its underlying table underneath it if you want to check the numbers.

One more wrinkle. WhisperTyping caps a recording at 300 seconds, and when I hit the cap mid-thought I just start another recording and carry on, so a handful of recordings are really halves of one message. That happens ten times more often now than it did in September, 0.3% of recordings then against about 3% now, which is itself a sign of the messages getting longer. Stitching those continuations back together barely moves anything though: the medians shift by a word or two and the per-day decline gets slightly steeper, from 26% to 30%, so the charts count recordings the way the tool recorded them and the published figures are the conservative ones.

Fewer recordings a day, more words a day

The first thing I wanted was simply how much I was talking. Counting recordings alone is misleading if the recordings change length, so here is both, per active day, meaning a day I actually dictated something. That is not a working day in the Monday to Friday sense. As the last chart shows, plenty of them are Saturdays.

recordings per active daywords per active day
0306090120150180SepOctNovDecJanFebMarAprMayJunJul2025202615285
Table view
An active day is a day with at least one recording. July 2026 covers 1 to 25 July only.
MonthActive daysRecordingsRecordings per dayWordsWords per day
2025-09301,91163.7237,7727,926
2025-10271,23545.7149,2395,527
2025-11291,98368.4242,9378,377
2025-12301,07235.7134,2324,474
2026-01291,64356.7254,5078,776
2026-02221,20054.5186,8188,492
2026-03251,74769.9275,51711,021
2026-042890732.4163,3805,835
2026-051251042.5104,8578,738
2026-06301,11937.3245,1378,171
2026-07241,09045.4239,6979,987
Both indexed so that 100 is the average of the first four months, September to December 2025. A single base month would import whatever was odd about that month; the four-month average does not. Raw values in the table.

Comparing the first four months against the last four, recordings per active day are down 26% and words per active day are up 24%. On a day I dictate at all, I now speak less often and say more each time. The other thing hiding in that table is the number of days itself: September had 30 active days, July had 24. The month totals fell mostly because there were fewer active days, not because the days themselves got lighter.

The prompts nearly doubled in length

My median message went from 89 words in September to 164 in July. I would have told you the opposite happened. I assumed better agents meant less hand holding and shorter prompts, and the data says I used the improvement to say more per instruction instead.

Two things are tangled inside that rise, and they are worth separating. A recording counts as prompt-brokering if it mentions the word prompt, and a small number always did, which is why the blue line runs all the way back to September. Up to December those recordings carried only 5 to 11% of my words. In January I changed how I work: rather than mostly talking straight to one agent, I dictated into a second LLM and had it write the prompt for the coding agent. The strip under the chart shows that change directly: brokering jumped from 7% of my words in December to 42% in January, stayed near 45% until June, then fell back in July when I moved to an agent I mostly talk to directly again. Because brokering recordings are structurally longer, that shift alone lifts the overall median, which is why the chart splits the two modes.

talking directly to an agentdictating a prompt for another agent
060120180240300265157 Jan 2026 share of each month's words that sit in brokering recordings 0%25%50%7%42%SepOctNovDecJanFebMarAprMayJunJul20252026
Table view
A recording counts as prompt-brokering if it contains the word prompt. Crude, but it separates the two modes well enough to see the effect.
MonthMedian words, directMedian words, prompt-brokeringBrokering share of words
2025-0988102.05.6%
2025-1076119.56.4%
2025-1183.014711.1%
2025-12861546.8%
2026-0183.017342.2%
2026-028217142.9%
2026-0392154.045.1%
2026-04109184.547.9%
2026-05126.0179.042.3%
2026-06133219.047.4%
2026-07156.5264.516.2%
Top: median words per recording, split by whether the recording mentions the word prompt. The blue line exists before January because the mode existed before January; it just carried a small share of the words. Bottom: that share, month by month, which is where the January change actually shows.

The interesting part is the amber line. Talking directly to an agent sat flat at 76 to 92 words for seven months, then started climbing in April: 109, 126, 133, 157. That is the genuine change, and it lines up with when agents got good enough to run unattended for long stretches. The better they got, the more I gave them per instruction. Both modes are still growing at the end of the data.

Eleven months, and my mouth does the same thing

One number refused to move. I have 332.7 hours of audio with exact durations, so I can work out how fast I actually speak. It is 113.1 words per minute across the whole corpus, and the monthly figure never leaves a band between 110.5 and 116.2, which is 5.2% wide.

050100150SepOctNovDecJanFebMarAprMayJunJul20252026113.7
Table view
MonthWords per minuteHours of audio
2025-09110.535.9
2025-10115.721.5
2025-11113.735.6
2025-12113.619.7
2026-01111.937.9
2026-02114.327.2
2026-03116.239.5
2026-04111.124.5
2026-05111.715.6
2026-06112.036.5
2026-07113.735.1
Words per minute by month, raw, on an axis starting at zero so the flatness is the real thing and not a cropped scale.

Four primary models, prompts nearly twice as long, a complete change in how the work is organised, and I am still talking at the same speed I was in September. Whatever prompt engineering is for me, it is not a change in delivery. It is just more words at the same pace.

The politeness collapse, and the word that survived it

This is my favourite thing in the data and the one I would never have predicted.

I used to say please to the machine, constantly. It runs at about 21 uses per 10,000 words for the first four months, falls off a cliff in January and February 2026, and settles around 8. Thank you went down harder still, from 3.9 per 10,000 words in September to 1.0 in July. I stopped being polite to it and I never noticed I had.

Sorry did not follow. 1,523 uses across the corpus, bouncing between 5.0 and 8.3 per 10,000 words with no real trend. By February please had fallen far enough that the two lines meet, and from there they travel together.

"please""sorry"
0510152025SepOctNovDecJanFebMarAprMayJunJul20252026 23.7 5.0 8.2 6.8
Table view
Literal string matches on the words please and sorry. Rates are counts divided by that month's total words, times 10,000.
Month"please" countper 10k words"sorry" countper 10k words
2025-0956323.681184.96
2025-1031521.11805.36
2025-1150820.912008.23
2025-1227320.34705.21
2026-0131312.31957.66
2026-021015.411146.1
2026-032218.022298.31
2026-041297.9905.51
2026-05797.53777.34
2026-0625110.241787.26
2026-071978.221626.76
Uses per 10,000 words, plotted raw. Comparing the mean of the first four monthly rates against the last four, please is at 39% of where it started and sorry is at 108%.

I think they were never the same behaviour even though they look like it. Please and thank you are deliberate courtesies aimed at something I half thought of as a person, and eleven months of familiarity wore them off the way it wears off with a new colleague. Sorry is not aimed at anything. It is a verbal tic, three quarters of it mid-sentence corrections like "the routing panel, sorry, the routing window", and you cannot wear that off because it was never a decision in the first place.

I bury the ask, and I did not know I did

When I dictate a long message the request almost never comes first. I set out where things are, walk through why, and only then say what I want. Across 3,823 messages the median request sits 83% of the way through, and 54% land in the final fifth.

0%10%20%30%40%17%37%0102030405060708090100 how far through the message the ask appears (%)
Table view
3,823 of the 7,575 messages over 100 words contained a detectable request. Position is the start of the last matching phrase as a share of message length.
Position in messageShare of messages
0 to 10%6.7%
10 to 20%4.1%
20 to 30%3.6%
30 to 40%4.2%
40 to 50%5.0%
50 to 60%5.9%
60 to 70%6.6%
70 to 80%9.7%
80 to 90%17.2%
90 to 100%37.0%
Position of the last request phrase as a share of message length, across the 3,823 of 7,575 messages over 100 words that contained one. The phrase list is things like "can you", "I want you to", "let me know", "have a look". Take the first match instead of the last and the median is 78%, so the shape holds either way.

11% do lead with the ask, and those are not one-liners, because the sample only contains messages over 100 words. They are the ones where I gave the instruction first and justified it afterwards. The rest of the time the ask waits until the case for it has been made.

What I find interesting is that an LLM does the exact opposite. An assistant has only ever been trained to answer. Somebody asks, it replies, and leading with the answer is correct in that situation. But I am not answering, I am starting something, and when you start something the ask has to earn its place first. I suspect that mismatch is a big part of why AI-drafted posts read as pushy when nobody intended them to be.

The model carousel

Every recording is timestamped, so I can see which model I was on by which name I say out loud. Four primaries in eleven months.

CopilotOpusGPTFable
051015SepOctNovDecJanFebMarAprMayJunJul2025202614.09.8 6.7
Table view
Mentions per 10,000 words. Copilot includes the hyphenated spelling. One October 2025 hit for Fable was the video game and has been removed.
MonthCopilotOpusGPTCodexFable
2025-093.790.01.930.80.0
2025-102.551.071.740.340.0
2025-112.720.120.330.080.0
2025-121.640.60.220.00.0
2026-012.00.91.220.240.0
2026-023.5911.672.190.640.0
2026-030.9813.971.340.040.0
2026-043.436.619.794.350.0
2026-052.770.04.013.530.0
2026-061.025.753.711.792.33
2026-070.083.251.080.046.68
Mentions per 10,000 words. Codex is in the table rather than the chart because it tracks GPT almost exactly. One October 2025 hit for Fable turned out to be the video game and was removed, which is a fair warning about matching model names by string.

Copilot holds steady at two to three mentions per 10,000 words right through to May, then falls off a cliff: two mentions in the whole of July. Opus arrives in February and owns February and March, peaking at 14.0. April and May are a GPT and Codex detour, with GPT at 9.8 in April and Opus at exactly zero mentions in May. Then Fable appears on 11 June and is at 6.7 within six weeks. I did not consciously decide any of those switches. I just followed whatever was doing the job at the time.

The other thing in there is what is not in there. Grok gets one mention in 2.26 million words, Kimi none. For all the tool shopping, I never really left the big labs.

There is no weekend

0%2%4%6%000204060810121416182022 hour of day, share of all 14,625 recordings
Table view
HourShareHourShare
00:003.06%12:006.68%
01:001.87%13:006.67%
02:001.43%14:006.55%
03:001.23%15:006.11%
04:001.03%16:005.78%
05:000.96%17:005.94%
06:001.63%18:005.43%
07:002.08%19:005.61%
08:002.54%20:005.11%
09:005.22%21:004.96%
10:005.57%22:004.75%
11:006.19%23:003.62%
Share of all 14,625 recordings by hour. Amber bars are 22:00 to 06:00. The peak is midday, and the decline through the evening is steady, but 3.1% of everything still lands in the midnight hour, more than lands at eight in the morning.

679 recordings land between 2am and 6am. Normalised per calendar day, Tuesday is my busiest day and Friday my quietest, and Saturday and Sunday sit in the middle of the pack rather than at the bottom. The weekend does not exist as a category. I was active on 292 of 334 days, with only four breaks longer than three days in the whole period, the longest being ten days in May.

The dataset also passes the sanity check I wanted from it. The word hackathon appears zero times in 2.26 million words until 12 February 2026, the day the EVE Frontier hackathon was announced, then jumps to 19.6 per 10,000 words that month. The event itself ran 11 to 31 March, so everything before that was planning: a staging repo, a local devnet, a lot of markdown documents. My three busiest days in the entire corpus, 3 March (193 recordings), 2 March (175) and 18 February (153), all land in that planning window rather than in the hackathon itself, which I think is because a planning document comes back in minutes where code takes an agent a long stretch, so the dictation loop spins much faster. If a spike that obvious had not landed where I knew it should, I would not trust anything else on this page.

Why any of this matters

I did not set out to measure myself. I was trying to build a writing style profile so that when I ask an LLM for a Discord post it does not come back sounding like a press release. The measurements were a side effect, and they turned out to be more useful than the thing I was after.

The bit I keep coming back to is that eleven months of daily use changed almost everything about how I work and almost nothing about me. The tools turned over four times. The prompts nearly doubled. The manners quietly evaporated. And underneath it all I am still talking at 113.1 words a minute, still explaining myself before I ask for anything, still calling myself a vibe coder in the same breath I used in week two.

If you dictate to agents and you have an archive sitting there, it is worth a look. I would be interested to know whether the politeness thing happens to everyone or whether that one is just me.

voice dictationai coding agentsvibe codingprompt engineeringllm workfloweve frontieref-mapdeveloper productivity