July 29th, 2026
(This blog post about AI is very long, nearly 6,000 words… but Google can summarize it for you!)
There’s a new scholarly paper available, published July 22, 2026, titled “Generative AI floods and dilutes the market for books.” The authors are Tuhin Chakrabarty, Xinyue Liu, Jane C. Ginsburg, and Paramveer Dhillon. It might have remained obscure, as academic papers often do, except that Alexandra Alter made it the focus of her July 28 New York Times article, “How A.I. Books Sneak Their Way Into Stores.” 
Since then it has received significant coverage — today (8/2/2026) I find about 100 articles and social media posts referring to the article. I’ve spoken to several of my publishing colleagues in the past few days; all of them have accepted the paper at face value, as did the respected and widely-followed AI commentator, Ethan Mollick. Who digs deep into a 30-page academic paper strewn with mathematical formulae?
In this case, I do.
And I do not accept the paper’s assertions. In this post I hope to explain why.
I see the paper as part of the wider AI witch hunt. When I requested comment on an earlier draft of this post Chakrabarty replied to me, “Sorry I don’t think I have the time to do that. I also think you are advocating for use of AI to write books & morally that’s not okay with me so I can’t engage further.” I don’t see AI use as a moral issue, and don’t think using a moral lens clarifies the complex issues involved.
Keep in mind that the article was posted as a “pre-print,” and is not peer-reviewed. People outside of scholarly publishing mostly don’t understand the significance of peer review. It’s a rigorous quality control process routinely applied before a scholarly paper is published in a respected journal. And so some parts of this paper are likely to be rewritten/restated before final publication.
Also, as I illustrate below, the paper is math-heavy. There are lots of dense descriptions and, for me, inscrutable formulae. I’m sure that all of this was necessary to advance the necessary arguments, but I found less clarity wading through the math, not more. It reads like it was written to withstand cross-examination in a courtroom.
While I’ve read the paper several times, it’s possible that my misunderstanding of some of the math will have led me to incorrect conclusions about the substance of a particular argument.
I want to create a context for this post, and the study.
On the one hand, the study is simply a mathematically-rich attempt to classify the percentage of recent self-published genre fiction books on Amazon that are, at least in part, AI-generated. And then to try to determine what the financial impact of these books is on authors that do not use AI in the text of their genre novels. Fair enough.
But, underlying the study are a range of contexts, most not explicitly stated. One is that AI-generated fiction is a bad thing, per se. Another is that books with AI-generated text reduce the sale of books that do not contain AI-generated text, and that the “good student” AI-free authors need to be protected from these books. Another is a belief that AI detection software can reliably identify text created using AI software. Another is that author fears about, and anger towards, AI, is a reasonable proposition that should be fundamentally supported through investigation and litigation.
When drafting this post I hadn’t paid attention to the credentials of the co-authors. George Walkey, in his Context Window newsletter, points out that one of the co-authors is Jane Ginsburg, “one of the leading scholars of copyright and IP.”
The other co-authors are Paramveer Dhillon, an associate professor at the School of Information Science, University of Michigan and the MIT Initiative on the Digital Economy and Xinyue Liu, “a first-year PhD student at Stony Brook University, advised by Tuhin Chakrabarty.”
Chakrabarty, Ginsburg and Dhillon co-authored another paper on AI and copyright, “Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers,” which demonstrated that “the median fine-tuning cost of $81 per author represents a 99.7% reduction versus typical writer compensation… thereby providing empirical evidence directly relevant to copyright’s fourth fair-use factor, the ‘effect upon the potential market or value’ of the source works.” Chakrabarty and Dhillon co-authored yet another paper “Can Good Writing Be Generative? Expert-Level AI Writing Emerges through Fine-Tuning on High Quality Books,” which raises “fundamental questions about the future of creative labor.”
Ginsburg was sole author of a 2025 paper, “AI inputs, fair use and the US Copyright Office Report,” that provides useful background detail for anyone trying to understand the legal issues that underlie all of these studies, particularly “the fourth fair use factor.”
No, I hadn’t realized the extent to which this group had already signaled their beliefs about the harm that AI is supposedly causing authors. They are all likely to be busy testifying in the months and years ahead about their much-discussed reporting.
Do You Remember the Wang?
Wang Laboratories was the company behind the Wang word processing computer system. Standalone systems sold for the equivalent of about $50,000 in today’s dollars. IEEE reports that by 1978 Wang was the largest worldwide supplier of word processing systems, with fifty thousand users. “In a few years 80% of the 2,000 largest US firms had bought Wang equipment. At one time, it was said, every secretary in America swore by Wang products.” By 1984, the company employed over 30,000, and the founder, An Wang, was worth $1.6 billion, ranking him “as the fifth richest American.”

A sailor uses a Wang word processor at the Atlantic Fleet Audio Visual Command office.
In that period word processing technology spread to other companies with similar devices, and very soon after to word processing software on personal computers. WordPerfect for the PC appeared in 1982. Microsoft Word was first released in October 1983.
That was the death-knell for companies like Wang. And it was the beginning of mass adoption of word processing as a tool for writers.
Many resisted, among them Isaac Asimov, Ray Bradbury and Alice Munro. Mordecai Richler stuck with his typewriter. Michael Ondaatje wrote in notebooks. Robertson Davies said, in 1987, “I don’t want a word processor. I process my own words.”
And now think back to the late ’80s and early ’90s, as internet use became widespread. The concerns quickly multiplied: job loss, scams, the death of traditional publishing, privacy loss, cyberporn, spam.
And then smartphones: challenges to literacy, social media saturation, mental health impact, and addiction to the screen.
Of course some of those problems pre-date the Internet and many pre-date smartphones.
But a lot of these concerns are strikingly real, and severe, and are with us today. Certainly both the writing and publishing industries have been heavily impacted by these technologies.
Yet somehow we cope.
Obviously AI is a far more complex and far-reaching technology than personal computers and smartphones. The potential threats from AI adoption are frightening.
Of course, a powerful technology like AI requires regulation, and that will come, as we see it beginning to develop in Europe. But, with open source AI software now ubiquitous, banning AI is not an option. We need to learn to cope.
Several Shades of Grey: The Use of AI in the Authoring and Publication of Books

Lawsuits everywhere!
AI in authoring and publishing faces fundamental concerns beyond all of the other concerns that people already have more generally about AI. First is the “original sin” of the American AI companies, purloining trillions of words, many of them covered by copyright, without compensation to the authors (and their publishers). This is the basis of 129 lawsuits. Even if the AI companies win in the courtroom, they may never reclaim the trust of writers and other artists.
The second is an apparent assault on the most fundamental aspect of writing, the creative process itself. The perceived threat is as chilling for authors as concerns about copyright, energy use and the AI armageddon. These are firmly-held beliefs; for many they are marked in stone. I might not endorse them, but would never declare them “wrong.”
The Other Voices
Not as often heard are the voices of authors and publishers who see AI at the core of writing and publishing moving forward. Hugh Howey wrote an eloquent piece this past week, on the one hand a cri de coeur for the future of writing as a creative endeavor, on the other hand an acceptance of the inevitability of change.
Back to the Paper at Hand
Chakrabarty’s research paper (I refer to it throughout this post as Chakrabarty’s, as he appears to have taken ownership of the final) is complex, thorough and well-documented. It seeks to address a major issue regarding the potential damages in the current litigation surrounding how AI companies ingested vast quantities of copyrighted text — it asks, directly, what is the cost to non-AI-using authors from the AI-generated fiction being added to Amazon?
What Does the Paper Say?
As noted, this is an academic paper. In its PDF format it is 31 pages long. Exported to Microsoft Word, it totals nearly 17,500 words (including all footnotes, sources and captions). It contains some intimidating mathematical formulas, such as this:

and this:

No conclusions are stated that are labelled as such, but in the abstract of the paper the authors note that “books for which we detected substantial AI text… reach commercial scale, winning a growing share of sales over time and taking more of the scarce top-rank positions once held by books with no detected AI text…. Generative AI can… reshape a creative market through scale rather than quality. Our results bear directly on the market-effect question at the center of the fair use defense to copyright infringement.”
OK, So What Does the Paper Actually Say?
Reduced, and oversimplified, this is how I read the paper: Let’s say a bookstore carries a total of 10 different fiction titles by self-published authors. Those will sell a certain number of copies, perhaps a few, perhaps a bunch. Now, what if the bookstore also starts carrying three self-published novels generated, at least in part, using AI, and adds those to the 10 books already on display? Some people will buy those three new books, perhaps because they don’t realize that they were produced using AI, which might affect their quality, or perhaps because they seem worth reading. As a result, those 10 original titles by self-published authors will sell fewer copies.
Publishing and bookselling are essentially zero-sum games. People don’t buy hundreds of books because millions are available. Nearly every book that sells has sold at the expense of another book that didn’t sell, but might have sold if it faced no competition. The publishing industry in the U.S. hasn’t grown significantly in decades. More books get published each year, for the same number of dollars. My sale gained is your sale lost. (A big footnote here about just how wrong that statement is in some of the statistical details. But the point still stands.)
And so, in that sense, Chakrabarty’s paper is accurate. New books on Amazon, whether created with AI or not, cannibalize the sale of existing books on Amazon. Oh, but my god, there are already so many books on Amazon, so very many. Do AI-generated books truly upset the apple cart?
Are AI Books “Flooding” the Market?

It was a very good year.
How can you flood a market that’s already flooded, already under water?
Way back in 1993 there were 7,721 new works of fiction published (or, in some cases, new editions of existing titles), out of a total of 104,124 new books, in hardcover and paperback. In 2025 there were roughly 475,000 new adult fiction titles published, from a total of over 4 million new titles and editions. That’s 71 times more new fiction titles and editions as were published in 1995. And, as Mike Shatzkin often reminded us, ebooks never go out of print, while dated print books generally find their way also to the used book market. Newly-published titles just get added to the vast catalog of existing titles, in multiple formats, available for purchase on Amazon and elsewhere.
And how big is the book catalog?
Google has estimated that, as of 2010, 129,864,880 books had been published. A Stanford Literary Lab study from 2016 raised that number to 189 million books (and estimated that perhaps 10% were fiction). WorldCat.org, a nonprofit that holds data from the collections of more than 10,000 libraries, estimated in 2022 that it had records for 405 million books (this would include multiple formats and languages) and 43 million ebooks.
Jim Milliot at Publishers Weekly offered an exclusive view of Bowker’s assessment of the “total number of books published in the U.S. in 2025 with ISBN numbers.” Which was 4,172,222. Self-published books numbered 3.5 million, up from 2.5 million in 2025. Traditionally-published books totaled over 640,000, six times the number of titles traditionally-published in 1993. (Keep in mind two things: 1. different editions of books, such as ebook and print, generally receive separate ISBN numbers, except that 2. self-published authors often don’t bother with ISBNs, which can be expensive, and are not required by Amazon [even though Amazon now offers them for free].)
And how many books are on Amazon? Up until about 2017 you could actually find that number right on Amazon. My last snapshot of the published number, in early 2017, found 47.7 million books (some with multiple editions). Ebook Friendly, reporting initially in 2021, claimed “over 10 million publications in the Kindle Store.” Derek Haines, at Just Publishing Advice, put the number at “over 12 million ebooks in the Kindle Store” in 2023.
Let’s go back a few years, to the pre-ChatGPT days. It’s safe to assume that few if any AI books published in 2021 were created with significant involvement from AI LLMs. And yet in the 5 years from 2017-2021, an average of 2.43 million books were published each year, of which, on average, nearly 300,000 were fiction. And all of these were stirred into a pot of at least ten of millions of existing titles.
Take any of the numbers, add in the data released since, and the totals are through the roof.
And yet some people are worried that AI is going to “dilute the market for books”? Amidst this vast clutter, it is already all but impossible to attract attention to a newly self-published book on Amazon. That is the problem here, not AI.
Amazon is (sort of) a Democracy
Not bad!
Everyone knows that you can buy reviews on Amazon (even though Amazon forbids it). People assume that review purchasing on Amazon is rampant. It probably is. But buying reviews costs money, perhaps $15 per review. At that price, it makes little sense to buy reviews for ebooks that mostly retail for less than 5-10 dollars. You need to buy a bunch — different sources suggest different ranges — but 3 or 4 reviews won’t move the needle. And nothing stops dissatisfied customers from leaving damaging bad reviews, amidst the good reviews, and it’s nearly impossible to have those removed.
Few people understand that the number of reviews, good ones in particular, also impact whether Amazon displays your book in search in the first place. Few or no reviews mean that a book remains largely invisible to browsers in the Amazon bookstore (unless you pay Amazon for advertising placement).
The result is that bad books tend to fail on Amazon. That’s only democratic. You see it all the time: one or two reviews, both bad, and a sales rank of less than 5,000,000. If you want an AI-generated book to succeed significantly on Amazon it has to be good; readers have to enjoy it, to tell others, and to be motivated to leave positive reviews.
Yes, of course, some bad books sell some copies, and they might never get reviewed one way or another. (I don’t think that Chakrabarty’s study was looking for those.) But that’s not how it mostly works. Bad books fail in the Amazon marketplace, AI or not. Millions of books fail each year, joining their kindred failures from years past.
How to Spot an AI Author
I’ve been working for several months now with Michael Fraiman on a project to try to identify how many of the bestselling books in Amazon medical categories are AI-generated. We suspect that it could be a large percentage. But we’re having trouble pinning that down.
Part of the challenge is in identifying bestsellers, because those change all the time — Amazon’s system is dynamic, close to real-time — a book that was a top bestseller at 9 this morning might be old news by noon (as you’ll see below).
Next is getting access to the full text to try to assess whether there’s AI content — you need to either buy the book or access it, if available, through Kindle Unlimited. Either way you need to break the Kindle DRM, which is legal for university researchers, but probably not for us civilians, and so we haven’t tried.
But perhaps the toughest challenge is one that Chakrabarty has reduced to a Pangram statistic — how can you assess if the book’s content is AI-generated — what does that even mean?
As far as we could tell in the couple of months we spent on the project (before setting it to one side, at least temporarily) there were few books containing 100% AI slop. Most authors, using the Kindle system, recognize that 100% AI slop is not a formula for wealth. Instead there were indicators that AI was probably used for some portion of the text, and perhaps for the structure of the book, and perhaps for the cover. But what did that tell us? OK, “some AI use.” Which means exactly what?
One thing we did see, consistently, was that trash books on Amazon, whether generated with AI, or just with poor intentions, usually featured make-believe authors. In most every case, the author bio was a generic, “Sally O’Connor lives with her husband and three cats on her family’s farm in North Carolina,” alongside a generic picture of a smiling 40-something female, and with no contact info nor a website mentioned. We would search high and low for an author fitting Sally O’Connor’s description and very soon realize what we has suspected: she was a fiction.
That became our benchmark: if the author is bogus, the book is suspect. Although it was not, in itself, proof of AI use. There are many ways to pull together a slew of words, call them a book, generate a sloppy cover, and upload the resulting mess to Amazon. (Remember the Wikipedia books scandal?) It’s a starting point; little more.
Let’s Look at Some Categories
Chakrabarty’s report focuses on eight genres, “roughly in proportion to each cluster’s share of new releases”:
- General Romance
- Speculative / Adventure Romance
- Fantasy / Supernatural / Horror
- General / Contemporary Fiction
- Mystery / Thriller / Crime
- Crime / Suspense Romance
- Sports Romance
- Science Fiction / Adventure / Dystopian
Arbitrarily, I’m going to choose “Crime/Suspense Romance” to examine more closely. I’ve selected that without first looking at any data on Amazon. I’ll go in fresh.
The first thing I find is that there is no category (or no longer a category) called “Crime/Suspense Romance.” There are seven categories under “Crime Fiction” but “Suspense Romance” isn’t one of them. There is instead a category under “Romance” called “Romantic Suspense,” so I choose that one.
(This points to an issue that Amazon observers know well, but is invisible or inconsequential to most. Amazon nominally supports the 5,500 or so official BISAC subject categories, of which an author is encouraged to choose their top three. But Amazon also reserves the right to do something that it does frequently, which is to re-assign books to other categories, or to rename existing categories, or to create new categories, and assign books to those.)
(A related issue, also mostly invisible, is that it’s up to the author to initially choose three subject categories for their book. Few authors understand how the 5,500-category BISAC system functions, and mistakenly, or sometimes deliberately, choose odd categories for their book. For example, one of the romantic suspense books on the top 25 list described below is listed under:
- Men, Women & Relationships Humor eBooks
- General Humorous Fiction
- Small Town & Rural Fiction
The hardcover is listed also under “Romantic Comedy” and “Contemporary Romance” and the paperback listed also under “Feel-Good Fiction.” None of the listings, including the audiobook, include the category “Romantic Suspense,” even though today the book holds a top 25 position on that list. So which is it then?)
Let’s head deeper into the category “Romance/Romantic Suspense.”
Chakrabarty’s paper refers often to the “Top-25 rank.” As the paper explains in section 4.6:
For each genre-week, we take each title’s median observed sales rank across its valid positive observations, order titles by that median within the genre-week, and define the Top-10 and Top-25 rank slots. We then aggregate slot composition by quarter and band, and for the genre heatmap we compute the share of the Top-25 slots held by books with substantial AI text in each genre-quarter cell. To measure top-rank churn, we compare quarter-to-quarter Top-25 title sets. Let 𝑇25 𝑔𝑞 denote the titles holding Top-25 rank slots in genre 𝑔 during quarter 𝑞. Retention of books with no AI text is the fraction of the prior quarter’s no-AI Top-25 titles that remain,
I really haven’t a clue what any of that means; my takeaway is that they put a focus on the top 25 sellers in each category. And so will I.
We’ll take a tour of the top 25 “Best Sellers in Romantic Suspense” as of 10:40 pm on August 1, 2026.
Nineteen of the authors have over 10,000 reviews. Eleven of the authors have garnered over 75,000 reviews for their books. All but one has other books available, generally many other books. Of those below 10,000 ratings, all six check in as actual humans, with credible online presences. Two-thirds (sixteen) of the authors’ books are available on Kindle Unlimited, nine are not.
Here, for your perusal, is my spreadsheet of the Top 25 Romantic Suspense books.
Can we take a chance and accuse 20% of these popular authors of adding 26% or more of AI-generated text to their bestselling books? And what would that mean; what would its significance be? A moral issue? For some. A sales or a reader satisfaction issue? Apparently not for most.
There’s another problem with using sales rank and bestseller lists on Amazon. Sales rank is a different beast than traditional bestseller lists. Those lists record total sales over a set period, usually one week. Amazon sales rank is recalculated roughly hourly, and so it is picture of sales results at a moment in time, not over the longer window. It’s common for a book to hop on and off a rank, and then back on and off again.
Sixteen hours after preparing the first list I went back onto Amazon to again capture the top 25 “Best Sellers in Romantic Suspense” and compare that to last night’s list. Only three of the books had the same bestseller position. Seven of the titles had disappeared, and seven new ones were included. These large shifts within short timeframes, alongside uneven subject classification, make Amazon sales rank murky.
It appears that Chakrabarty and his team attempt to compensate by taking multiple readings over time. (For example: “In each genre-week, sampled titles are ranked by median Amazon sales rank and the 25 highest-ranked sampled titles are retained; weekly slots are then aggregated to quarters.”) But, fundamentally, Amazon sales rank is an unreliable statistical indicator.
Is 26% a Substantial Amount of AI Text?
One of the notable features of the study is how it stratifies AI content. The authors classified “each book by the share of its text detected [by Pangram software] as AI generated, grouping titles into those with no AI text, light AI text (≤25%), and substantial AI text (>25%).” The least of my problems with this claim is the percentage. 26% or more is at least a chunk, perhaps a bunch, maybe a lot. And >25% is inclusive of 99% AI text, even 100% AI text, which surely is substantial.
It’s at the lower percentages that I’m uncomfortable. Let’s think through a scenario where an author generates about a quarter of their work with AI. Maybe a third. Perhaps even half. Let’s think where that text might appear in the work. What’s unlikely is that it will appear in an unadulterated chunk. More likely it will be woven into the larger work, a sentence here, a paragraph there, and another sentence or two a page later.
That takes work. And what takes even more work is the 74% or the 66% or the 50% of the text that’s not AI-generated. Assuming it’s not just plagiarized, the author has to do the same thing that non-AI-using authors must do: write. This is obvious of course, but I highlight it because of the tremendous bias that people have about AI writing. Authors using AI, for the most part, still have to work hard to write their book, if they’re going to succeed. And they need to make the tough time-consuming decisions that every self-published author must make: will I hire an editor, or a copy editor, or a proofreader? Will I hire a cover designer? Will I add illustrations? Do I want a well-designed interior for the book? Will I advertise the book? Promote it on social media? And so on.
Does AI Write Books?
Keep in mind that there is no machine named ‘AI’ that, unbidden, writes and publishes books on Amazon. Humans do the work. Yes, in many cases humans “cheat” by using AI tools to shift some of the difficult creative work of writing onto the machine. And no one likes a cheater. But, increasingly, serious authors are using AI tools in their writing, whether to generate some text, to analyze draft text, or, in the case of Nobel Prize winning author Olga Tokarczuk, for research. We need to get used to it. As Hugh Howey notes, this is how authors are going to work going forward. They will use tools that allow them to be better at their craft, and to succeed in the marketplace. Slop has always been slop, and always will be. The good stuff, the books people most enjoy, will, over time, rise to the top, as they always have. That’s how publishing works.
Amazon’s Strong Stance Against AI Books (kidding)
As Alexandra Alter reports in her piece about Chakrabarty’s article, “(Amazon) asks self-published authors to disclose to Amazon whether or not they used A.I. to create their books [but doesn’t enforce the rule], and limits the number of books each author can upload to the site to 10 per week.” 10 book per week! 520 per year! It really is time for people to loudly ridicule Amazon’s efforts around quality control for books, and not to take those efforts the least bit seriously. Amazon is, unmistakably, part of the problem, not part of the solution.
The Pangram Challenge
I’ve worked with Pangram software for several years now, and know one of the founders, Max Spero, collegially. I’ve written before that I have “a lot of faith in the seriousness of his technical approach.” Chakrabarty’s study relies heavily on Pangram software for its analysis of AI content. I would probably have done the same for a study such as this. But I don’t find a note in the paper that reminds the reader that Pangram, like other AI detection tools, has encountered lots of well-documented instances of false positives. Indeed, one of the books analyzed in the paper, Daggermouth, has led to controversy, as the author vehemently denies AI use. (This was first identified in a thoughtful article in The Atlantic, from “an underlying data set that the (academic paper) authors shared,” perhaps not a wise decision.)
While I do agree with the consensus opinion that Pangram is the best there is, the repercussions from a false positive can be drastic. That demands care.
How Do You Lawfully Obtain 14,000 Books? With Kindle Unlimited (KU).
The paper’s analysis of 14,419 tomes states that “all books analyzed were lawfully acquired through purchase, library lending, or directly from authors.” It stretches credulity to imagine that many of the authors provided free copies of their text directly to researchers trying to identify if the book was written with AI. And libraries carry very few books from self-published authors. That would leave “purchase.”
Let’s say, arbitrarily, that the researchers managed to get 419 ebooks directly from authors or libraries (could be more, could be less). That would leave some 14,000 to obtain with cash (could be more, could be less). Fiction tends to hover in the $2.99 to 6.99 zone, and so it could cost between $40,000 to $100,000! Wow. The other way to obtain them, for far less cash, and completely legally, is via Kindle Unlimited (KU).
Kindle Unlimited costs subscribers $11.99 monthly to access a huge selection of ebooks, audiobooks and magazines (“over 5 million digital titles“). You don’t purchase books on KU, you download and “read” them — they also remain available for customers to buy outright on Amazon. For the subscriber it can be a great deal. Voracious readers save lots of money, and encounter new authors at a low cost.
To enroll in KU, an author first enrolls in KDP Select, a free 90-day program for Kindle ebooks, which makes the book eligible for free book promotions and Kindle Countdown Deals (KCD). All books enrolled in KDP Select automatically become part of KU.
But many authors ignore KU, in part because it’s a very different business model (paying per page read, from a fund, rather than per book, as a percentage of the retail price). Some reviewers argue that the books on KU are of lower quality generally, “with the bulk of available books having a ‘B- and C-list kind of feel to them’.” It’s reasonable to assume that an author of a slop AI book would nearly always list the book on KU. Why would they not?
KU provides authors with a lot of money, an average of about $62 million/month; a total of $744 million in the last full year. On a per page basis, though, it’s less than half a cent, or $0.004615. Two-hundred-and-fifty pages of compelling prose earns the author $1.15. Ten pages read earns 4-1/2 cents.
Traditional publishers eschew KU because mostly they’re not in favor of a subscription model, not in favor of the company their books would keep, and not in favor of the sales restrictions that Amazon enforces. (I just checked the New York Times bestsellers on Amazon and the only titles available via KU were Matt Dinniman’s “LitRPG” books, featuring Dungeon Crawler Carl and his ex-girlfriend’s fluffy cat, Princess Donut.) KU subscribers know that those bestselling books won’t be available with their subscriptions, but tons of genre fiction will be.
Enrolling in KU also prevents the author from offering their book on other platforms like Apple, Barnes & Noble, Kobo — even the author’s own website — while the book remains enrolled in KU.
One statistical distortion caused by KU is that a book’s popularity on KU often differs from its overall sales rank. Individual copies of a book can sell poorly, but the book still can do very well with KU “reads.” (Lots of self-published genre authors report earning 80% or more of their income through KU.) For sales rank purposes, each download is treated by Amazon as the equivalent of a purchase. And so books that are popular on KU tend to rank higher overall than similar books that aren’t available on KU.
But the author earns income on KU solely on the number of actual pages of their book that are read. Just getting downloaded provides no income. The complex formula is well-described here. There is no method available to estimate the page reads for a book, nor the KU income. Chakrabarty writes, “We measure Kindle Unlimited as whether a title is available on the service, not as how much of it readers actually read. The panel does not tell us whether a given unit is a Kindle Unlimited borrow, a page read allocation, or an ordinary purchase.”
An interesting aspect of KU is that a book’s income there may relate far more closely to quality than it does under royalty systems. If a reader downloads a low-quality AI-generated book on KU, starts to read it, and recognizes the low quality, they will stop reading and move onto another book. The author will earn an insignificant amount of money. On the other hand, if a reader buys the same book, the author receives their full royalty (unless the reader goes to the trouble of returning the book and seeking a refund).
An AI-generated book on KU will only earn significant page revenue if readers find it to be of quality sufficient to match the genre books they are used to reading on the platform.
With these factors in mind, the prevalence of Kindle Unlimited titles in this study appears to be a distorting influence. First, AI-generated books are more likely to appear on Kindle Unlimited than they are more broadly on the Amazon Kindle platform. Second, there is no clear method available to estimate a book’s actual KU income.
7/30/2026 — I heard from Tuhin Chakrabarty on Bluesky that they used a KU subscription. “Most books are Kindle Unlimited so you can access full text at 11$/month” he wrote. And, as mentioned, that’s a perfectly legal way to obtain books. But it’s not the same as “purchasing” thousands of titles.
Machines Have Been Used to Analyze Writing for Centuries

The real story of AI
I’ll head off path here and heartily recommend Denis Yi Tenen’s magical work, Literary Theory for Robots: How Computers Learned to Write. Among many other gems, the book takes us back to the 14th century scholar, Ibn Khaldun, and his ‘zairajah,’ whereby “a learned soothsayer would write down a question, converting its letters into numbers. These would then be transposed back into letters, by consulting a number of intricate charts, according to ‘well-known rules’ and ‘familiar procedures.’ Several such computational ‘cycles’ would produce a shortened string of letters, which could finally be expanded into a rhymed answer. With proper training, the zairajah could obtain the ‘knowledge of the unknown, from the known,’ Ibn Khaldun explained, giving the intellect an ‘added power of analogical reasoning.'”
Later we visit with Athanasius Kircher (born around 1601), who devised a “Mathematical Organ…. Made of polished, ‘artfully painted’ wood, the Organ resembled a large box. Opening its lid revealed a row of labeled wooden slates, filed vertically: four columns of fifteen narrow slates, four columns of seven wider ones, and four columns of five of the largest pieces. When pulled out, each of the slates contained a string of letters. By consulting the included booklets— which he called, wait for it, applications—any combination of the planks cohered into a complete, harmonious composition. Cleverly, depending on which manual was consulted, the same organ could be used to compose music, write poetry or secret messages, and even do advanced math, such as reckoning the Easter calendar.”
Do these sound just a bit like AI today?
We reckon with AI without a view to the past, and I think that is a mistake, a mistake that underlies Chakrabarty’s study. AI books are not suddenly flooding Amazon. The sluice gates were already wide open. No doubt each new book added to Amazon serves, minutely, to dilute the market for the other 50 million plus books on sale. But that is not an AI problem. It’s a publishing industry problem, one that we have danced around for far too long.
