Crosspost: Who wrote this? Measuring the human contribution when writing with AI
Prof Tzachi Zach discusses his rewarding collaboration with Claude that both improved the quality of the article about his SEC comment letter tracker and made producing it more efficient.
This is a guest article from Tzachi Zach, a Professor of Accounting at the Fisher College of Business, Ohio State University. He has been the department chair since August 2025. He has also served as the Academic Director of the OSU Master of Accounting (MAcc) Program. Before Ohio State, Tzachi was an Assistant Professor of Accounting at Washington University in St. Louis for six years. He holds a PhD in Business Administration (Accounting) from the Simon School of Business at the University of Rochester and obtained his B.A. and M.A. from Tel Aviv University in Israel.
The “veteran financial journalist” outside editor Professor Zach refers to is me. I learned a lot by participating in the process.
In a previous article I wrote about my SEC comment-letter tracker, a behind-the-scenes look at how we[1] read more than 8,000 comment letters on the SEC’s semiannual reporting proposal. Claude did the mechanical work, writing (actually typing) most of the words in it. But the critical thinking was all me, as I will explain later. It was a process that I believe was a rewarding collaboration that both improved the article’s quality and made producing it far more efficient.
We kept the receipts: twenty-six “frozen” versions after the initial rough draft, a log of every handoff, and the text of the full conversation between us. Those receipts allowed me to ask a question that I think matters far beyond my own writing: when an AI “types” most of a document, how do we measure what portion the human writer influences or contributes?
This essay describes how the collaboration worked, shows a few examples of our back and forth, and then turns to the measurement question. That’s essential because “the AI wrote it” and “I wrote it with AI” are very different claims, and right now there is no accepted way to tell them apart.
How the back and forth worked
The mechanics were simple. One live file held the essay. Every time authorship changed hands, we “froze” a numbered snapshot before the next edit, so the document’s history became diff-able, allowing the differences to be isolated and analyzed the same way software developers track changes to code. (My next task is to Git it).[2]
Claude wrote the first draft from the project’s own records after I gave it the topic, the angles I wanted explored, and my voice, or personal writing style, rules. Then I revised. Then Claude revised. Twenty-six versions later, we had a pre-final text ready for the outside editor’s review.
My revisions took two forms, and the distinction turned out to be the heart of the whole exercise. Sometimes I engaged the text directly, almost always by dictation (via Wisprflow), rarely by actually typing; almost all of my commands to Claude were oral. I experimented with the text: tried a phrasing to hear how it sounded, moved a section, cut what did not sound like me, reframed an argument, and read it back. My contribution ran through ideas and direction far more than through composed sentences.
Sometimes I left instructions for Claude inside the draft itself, in capital letters, at the exact spot where I wanted the work done. For example:
“PROVIDE LINKS TO SOME EXAMPLES; GIVE AN ACTUAL NUMERIC EXAMPLE FROM THE DATA; THIS IS A BIT EXAGGERATED BECAUSE I DID NOT LITERALLY WRITE.”
The caps convention emerged naturally, and we codified it as a rule. Because the snapshots are frozen, every instruction is preserved verbatim at the version where I gave it. The draft carries its own paper trail.
Three illustrative examples of human-AI interactions
The first example started as an aside. In my first revision, I wrote that Claude’s ability to de-duplicate letters is extraordinary. That was pure enthusiasm. I asked for nothing. Claude turned the enthusiasm into an anecdote about catching duplicates. Then, after my reminder based on my recollections from a few weeks earlier, Claude went looking for a concrete case that would make the anecdote checkable.
It found two near-identical letters, in a corpus that was already up to more than eight thousand. (I remembered it was in the 200s and so I directed Claude there. ) The two letters were filed the same day by Jennifer and Mark Simpson, where one called the cost savings “miniscule” (sic) and the other a “nothing burger.”
A human could have found that pair after some significant effort, with scanning the folder, scratching head a few times, but I suspect it would take more than a minute. Claude did it quickly and efficiently, leaving me focused on thinking what else to get into the article, instead of being upset that I could not find those vaguely recalled letters. In summary, my initial contribution of one sentence of excitement turned into the reader getting an illustrative example through a verifiable fact with links.
The second example ran in the opposite direction, from human suspicion to a document-wide standard. The earlier “how I did it” essay used three letter writers listed on the docket as CEO and CFO to illustrate a classification problem. On my fifth read, I flagged something: “THESE LETTERS DO NOT SAY CEO.” Claude verified it: the titles come from the docket’s submission form and appear nowhere in the letters, so a reader clicking through would never see them.
Claude rewrote the example around exactly that gap, which was a more interesting story! Then, unasked, it applied the same test to every other credentialed person quoted in the essay, found two more who failed it, and replaced them with writers who state the credential in their own letters. One human-initiated flag became a fact-checking standard for the entire piece.
The third example shows who decides what gets in the essay. I left a one-line instruction asking for an actual kappa number, the statistic that measures agreement between raters. Claude came back with more than a number. It had found three different kappas living in the project’s files, explained which question each one answers, and recommended against the figure sitting in the draft, because it was the least informative of the three. I accepted the better numbers and then overrode the placement: the discussion belonged in the paragraph about human raters, where the comparison between expert readers and the machine actually lives.
The final passage is human-initiated, AI-surfaced, AI-drafted, and human-positioned. All four verbs matter to the final result.
Two ways to count
The initial draft, more of a rough outline, was written by AI, although it was directed by humans. But let’s assume for a minute that it was completely AI-driven. How much of the final essay is human-directed? What is the actual human contribution? How different is the last draft from the first?
These are important questions that researchers and educators like me are very interested in:
Is attribution important when an AI collaboration exists?
How do we evaluate student outputs?
How do we evaluate the interaction between students and AI?
The first measure one could reach for is textual. Walk through all twenty-six versions, tag every surviving word with the author who “typed” it, and count. By that measure, the final essay is about 24 percent of Claude’s first draft, 48 percent of Claude’s later passes, 26 percent of my own edits, and 2 percent from an outside editor. Roughly three quarters AI-typed. It’s a defensible number, computed mechanically.
But, I think this rough calculation badly misleads the reader.
That’s because I do not contribute by “typing.” I direct the ideas. One passage in the original “how I did it” tracker essay pairs two accounting scholars who filed dueling comment letters, Thomas Bourveau and Elizabeth Demers, on whether quarterly reports create information spillovers that are useful for other firms. Claude wrote every word of that passage, located the letters, and linked them. Yes, the passage exists because my instruction named both authors. Counting those words as AI influence gets the causality backward.
So, we built a second measure. Call it the provenance measure: instead of asking who typed each word, it codes each change by who caused it. Take every substantive change across the twenty-six versions, eighty-three of them by our count, and code each one by who caused it: changes I directed and Claude executed, defects I flagged and Claude fixed, edits I made myself, and changes Claude initiated on its own. The result: 87 percent of the changes were human-caused, 84 percent by me.
The AI-initiated category was small, eleven episodes, and every one of them was an act of verification, retrieval, or plumbing: a premise corrected, an audit extended, a stale graphic held back, a wrong count softened. Claude never autonomously added an idea or an argument (although it certainly could, in principle; and I have seen that in action, too).
The draft also passed through a second pair of human hands. Late in the process an outside editor, a veteran financial journalist, marked up the piece with tracked changes. Claude merged them, but it held three of her edits because they needed my call. Claude, per my instructions, returned her margin comments to me as questions instead of applying them quietly. For one round, the AI sat between two humans as a fact-checker. Her surviving words come to about 2 percent, but that number understates her influence for the same reason my 26 understates mine.
In sum, the two numbers answer two different questions. Claude typed roughly three quarters of the words. I caused 87 percent of the changes. Both are accurate, and I would not trust either one alone to say who wrote this essay.
The components
After we walk through some examples, I think it is worthwhile to reflect on AI-human collaboration, attribution, and measuring the role of human-AI interactions, and by extension, and most importantly to educators, student-AI interaction. My little exercise suggests the factors any serious measure of human influence on an AI document should capture.
Who initiates. The single most informative datum about a change is whether a human asked for it. That requires provenance records; without the frozen versions, who initiated is unrecoverable after the fact.
Who verifies. Some of the highest-value human contributions subtract words rather than add them. My CEO flag prevented a published error. A measure that only counts additions will miss an important part of the contribution.
Who decides framing and placement. The kappa episode shows that the same facts, in a different paragraph, tell a different story. Positioning is authorship.
Direction density and leverage. By leverage I mean where the human’s time goes: it shifts away from mechanics like typing and toward thought, and the AI multiplies what an hour of thought produces. It is measurable, and we tried. My caps instructions across all my revisions total about 300 words, dictated in minutes, and they replaced something like 10 to 20 hours of solo retrieval work. Claude’s revision passes left about 1,700 words standing in the final text, roughly 6 words of essay for every word of instruction. Words of instruction per words of output is a crude ratio, and it may not be the right measure, but I suspect it separates directed writing from delegated writing better than any word count.
Survival. Only about 41 percent of the AI’s first draft survived to the final version. The survival rate of AI text under human revision is itself a measure of how much the human is actually reading, judging, and discarding.
Voice. Whose sentences do they sound like? I maintained an explicit voice profile, my banned words and structures, and Claude audited every draft against it. The document reads as mine because that constraint was enforced, and enforcement is a human act even when the checking is automated.
None of these factors is sufficient alone, and I don’t yet know the right way to weight them. But I am confident about the direction: decision provenance is the primary measure, and it can only be computed if the collaboration leaves a trail. Freeze the versions. Keep the instructions in the draft. Save the conversation. I tried a somewhat initial version of this while grading student AI collaboration this past semester. I think it worked pretty well, although I didn’t quantify it at the time. I hope to write about it at a later date.
(The provenance measure returns in the Measurement section below, where it meets the AI detectors.)
Why does all of this matter?
Universities, academic journals, and courts (and many others) are all reacting and improvising rules about AI-assisted writing. The universities are naturally where my interest is. At universities two legitimate AI-related goals collide. One goal is to keep AI from stealing students’ brains: an MIT Media Lab study (still a preprint) measured the weakest brain engagement in students who wrote essays with ChatGPT, and they remembered the least of their own text (TIME covered it), while a 666-person study found that heavier AI reliance went with lower critical-thinking scores, particularly affecting younger participants. The other goal is to produce graduates who can work effectively with AI, and it is just as pressing: my own university now requires AI fluency of every undergraduate, and two thirds of business leaders say they would not hire someone without AI skills.
An AI ban serves the first goal and starves the second, and most university professors, and administrators, are asked to serve both at once. The enforcement record shows the strain: universities adopted AI detectors and then began walking away from them, with Vanderbilt disabling Turnitin’s detector after calculating that even a 1 percent false-positive rate would wrongly flag about 750 of its students’ papers a year. NPR reporting, three years into the ChatGPT era shows that professors are still enforcing homemade rules that do not always agree.
The academic journals are improvising too. Science and Nature took opposite positions in the same month and then traded places: Science banned AI-generated text outright in January 2023 and reversed itself that November, while Nature allowed disclosed use from the start. Shiva Rajgopal and Robert Eccles describe the wider wrestling in accounting and finance academic publishing.
The courts, too, are learning by sanction: the $5,000 fine in Mata v. Avianca for six fabricated citations has grown into a running database of more than 1,700 court decisions involving AI hallucinations as of this writing. The Fifth Circuit proposed a rule requiring lawyers to certify their AI use, then dropped it a few months later.
Some of this improvisation is understandable because the thing being regulated continued to change under the regulators’ feet. The models of 2023 are not the models of 2026. When those first policies were written, the state of the art had just made a startling jump: in March 2023, OpenAI reported that its new GPT-4 scored around the top 10 percent of test takers on a simulated bar exam, while GPT-3.5, the model behind the original ChatGPT released only four months earlier, had scored around the bottom 10 percent.
Since then, Stanford’s AI Index documents the cost of 2022-level output falling more than 280-fold in two years, and METR finds that the length of tasks AI can complete on its own has doubled roughly every seven months. The machine-learning conferences describe the same arc in policy form: ICML’s 2023 ban called itself provisional (”this policy may evolve”), and by ICLR 2026 the rule had become disclosure plus author responsibility, citing “the changing landscape of LLM usage.” A rule written in January 2023 was written for a different machine.
Measurement
Assessing AI authorship and contribution is not merely a text-scanning exercise. We already have the alternative in hand: the provenance measure from Two ways to count, the eighty-three coded junctions that put the tracker essay at roughly 72 percent AI-typed and 87 percent human-caused. Yet the rules and the tools treat assessment as a text scan, as if the share of machine-typed words represented the share of machine thinking. The detectors have improved dramatically since 2023, but are supposedly detecting an answer to the same question: Does the text scan as AI-generated?
Early on, detection meant scoring statistical properties of the words: GPTZero explained its method through perplexity and burstiness, and even OpenAI, which knew its own models best, shut down its detector because it could not tell the difference reliably. Today, GPTZero describes an end-to-end deep-learning classifier, Pangram trains a transformer on large corpora of AI text and claims a false-positive rate of one in ten thousand, and an independent University of Chicago audit found Pangram’s error rates near zero on longer passages while a 2023-generation tool misclassified up to three quarters of human text as AI.
The engineering improved; the question did not. Each of these systems reads the finished words and only the finished words. Tellingly, GPTZero sells writing-process verification as a separate product, with edit-history replay and typing-pattern analysis. Its own advice to educators facing a positive detection is to ask the student for drafts and revision histories. Even the detection industry treats the trail as the ground truth, because the detector itself cannot see how a document came to be. A perfect detector of machine-typed text would still measure the wrong thing in a collaboration where the machine types what the human causes.
So, I tried a small, informal experiment: one document, for illustration only. I ran the earlier tracker essay, whose provenance we know down to the word, through two detectors. GPTZero read it as “likely a mix of AI and human.” Pangram read it as “Mostly Human, AI Detected,” with 18 percent AI content. Both are reasonable readings of the words, and a single article says nothing about either tool’s accuracy in general, or about which is better. The point is different.
Recall from Two ways to count that the words of that essay are about three-quarters machine-typed and 87 percent human-caused, and no reading of the finished text, however good, can recover that pair, partly because AI-typed words audited into my voice read as mine. The detectors answer the question they were built for: does this read as AI? The question here is a different one: who caused it? Only the trail answers that.
The provenance measure asks the causal question directly: who caused each change, coded junction by junction the way we did in Two ways to count. And it dissolves the university tension I discussed previously: grading the junctions a student causes in an AI-drafted text, the directives, the flags, the corrections, and you are grading exactly the critical thinking the first goal wants to protect, while teaching the collaboration the second goal demands. Some educators are circling the same move away from the finished artifact and toward the interaction, proposing that faculty review students’ AI transcripts: “the conversation itself becomes the artifact of assessment.”
The tracker essay was the pilot. We measured the collaboration on that document after the fact: the piece that emerged is better than what either of us would have produced alone, and I can prove where every one of its parts came from. That proof took almost no extra effort at the time, just the discipline of snapshots. (The essay you are reading is getting the same treatment, snapshot by snapshot, as I write it.)
What if that discipline becomes the norm (and is embedded and automated, by, ahem…AI)? Because the question that matters is who caused the document to say what it says, and only the audit trail can answer it.
When I say we, I mean Claude, the AI assistant I work with, and myself. The work was a collaborative effort, and this essay is about exactly what that means.↩︎
For the non-programmers: Git is the free version-control system software developers use to record every change to their files, who made it, and when. Each saved state is a “commit,” and the full history can be replayed or compared at any point. Our frozen snapshots do the same job by hand; Git would do it automatically.↩︎
© Francine McKenna, The Digging Company LLC, 2026








