I love this conversation. Below is what I think and I talk in general, I haven’t looked at the particular PR mentioned.
I think adding a skill around writing docs/comments and actually settling on some directions about what we aim for as a community, so new people can also learn quickly, not just by vibes, reading around - sounds great. Thanks David, for bringing it up. “My 2 cents would be, when we review comments/documentation, the questions we ask ourselves should be: 1. Is it correct? 2. Is it clear? 3. Is it concise? And any other good qualities we have been looking for in documentation since pre-AI time. But not “does it look like AI-generated?” If the AI-generated doc is in good quality, what’s the problem?” I agree on this. I catch myself cutting/editing a lot of AI-generated docs/comments for good reasons stated here. But also, as a non-native speaker, I love how AI can quickly improve my writing, my grammar. And I wanted to mention - sometimes I ask it to improve what I wrote, then it feels a bit different style than the one I would use, but it is still accurate and makes sense - then I leave it as is but I also stress even more for others that AI helped, especially with the writing, but I verified it and I fixed it. It is different but verified and good, so I’d say it serves the needs of documenting more, which Chris mentioned we were missing a lot in time. There is also the psychological element - we all don’t want to allow slop, so the moment we see something that doesn’t sound exactly as the person who submitted it, many times a red lamp turns on, and we catch ourselves wondering how much the person who submitted it looked into it etc., but that doesn’t mean the submission is immediately bad. *This email was fully Ekaterina-written. :-) (I am sure there are some grammar mistakes to prove it)* On Sun, 23 Aug 2026 at 15:17, Jane H <[email protected]> wrote: > Thank you for bringing this up! > > My 2 cents would be, when we review comments/documentation, the questions > we ask ourselves should be: > 1. Is it correct? > 2. Is it clear? > 3. Is it concise? > And any other good qualities we have been looking for in documentation > since pre-AI time. > > But not “does it look like AI-generated?” If the AI-generated doc is in > good quality, what’s the problem? > > So, I agree that we should use the word API instead of “the load bearing > seam”, cuz it’s not clear. I also agree that we should get rid of the > verbose comments, cuz they are not concise. But I don’t think we need to > bother to make it looks human-written, like getting rid of the dash - as > long as it’s still correct, clear, and concise. > > Jane > > > > > On Sun, Aug 23, 2026 at 11:45 Chris Lohfink <[email protected]> wrote: > >> I guess I have a bit of the opposite view. After years of no comments >> and vague documentation we have correct albeit verbose PRs coming up. While >> I think it's perfectly acceptable to create PRs or make recommendations >> that do things like make comments more readable, I am -1 to the idea of >> restricting or having requirements verbiage to being a nebulous "better". I >> think that also leads to a path that can discriminate against non native >> speakers for difficult to understand English in their commits. >> >> Chris >> >> On Sun, Aug 23, 2026 at 1:35 PM David Capwell <[email protected]> wrote: >> >>> Maybe a native speaker can quickly digest and filter out all witticisms, >>> but for folks like me it might be challenging. >>> >>> >>> We can’t… there is a reason that 2 skill patterns have taken hold: speak >>> to me like i’m 5, and asd-ste100… in short it’s word vomit that no one >>> understands. What’s worse is that trying to understand it confuses humans >>> and Claude; leading to worse outputs than if the comments were just thrown >>> away. >>> >>> >>> Sent from my iPhone >>> >>> On Aug 23, 2026, at 6:23 AM, Alex Petrov <[email protected]> wrote: >>> >>> >>> Thank you for bringing the subject of LLM-generated prose. I'd like to >>> first mention that I only had a cursory glance at 21452/21462, and my >>> comment is about LLM-generated prose in general rather than this patch in >>> particular. >>> >>> My recent impression after reading an LLM-generated README in an OSS >>> project (albeit not in Cassandra) was the feeling that I have started >>> reading mid-paragraph, and the writing itself was witty and confusing, >>> which ultimately lead me to closing the browser window rather than >>> bothering to digest that text. >>> >>> I tend to remove all LLM-generated comments from my own code. For me, >>> clarity goes beyond comments in code: any artefact that another human is >>> going to read (for example, a code-review post, or bug >>> analysis/investigation), needs to pass a readability bar. Having that said, >>> I hate to admit that my own writing style might have also changed for worse >>> over the course of this year. >>> >>> An additional reason (besides consistency with our existing comments and >>> documentation) for having simpler language in documentation and code is the >>> fact that for many of us English is not a first (and for some, not even a >>> second) language. Maybe a native speaker can quickly digest and filter out >>> all witticisms, but for folks like me it might be challenging. >>> >>> +1 for simplifying the language for both AI- and human generated prose. >>> >>> >>> On Sun, Aug 23, 2026, at 2:01 AM, [email protected] wrote: >>> >>> Hi all, >>> >>> Anthropic’s current models are famous for generating overwrought >>> metaphorical constructions like “the load-bearing seam” when referring to >>> something as simple as an interface. r/ClaudeAI has dubbed this manner of >>> speaking “Claudish.” Many users (including myself) have elaborate user >>> prompts that try to tame the model, while others go as far as passing >>> Opus/Fable-generated output through a competitor’s model to untangle it. >>> >>> Like many, I find reading Claudish grating and artificial - like the >>> taste of a Sweet ’N Low packet (aspartame), or listening to a 48kbps MP3 >>> dominated by compression artifacts. >>> >>> I’d like to start a discussion about project norms regarding >>> model-generated comments and documentation in our codebase, largely >>> prompted by the merge of CASSANDRA-21462 (5b34068). >>> >>> I open with my gratitude for work to validate and harden cursor-based >>> compaction. My local measurements land it between 1.7 - 2.4x the throughput >>> of legacy iterator-based compaction – a stunning improvement that will make >>> Cassandra faster and more stable. I also appreciate the focus on >>> correctness and validation in this work, as it surfaced and resolved >>> several serious issues. >>> >>> The concern it prompts for me is that the commit marks the first >>> introduction of Claudish into the codebase, and quite a lot of it. >>> >>> Examples in the first 1/3 of the patch include: >>> >>> – The zero case is load-bearing rather than an optimisation >>> – The seam is EVENT-shaped because {@link UnfilteredDescriptor}s are >>> transient >>> – The row-side analogue of cell reconciliation, is load-bearing in the >>> cursor's row >>> – A decoder defect is as likely as on-disk damage here, but that's the >>> same ambiguity >>> – The mirror has already drifted from the upstream serializer once >>> _ The precondition is asserted rather than assumed >>> – The cursor compaction path and the reference path reach one decision. >>> They did not always: the cursor carried a hand-mirrored copy >>> >>> The linguistic style of Anthropic’s models is sharply out of step with >>> comments in Cassandra’s codebase. Our comments are concise, flat, and >>> matter-of-fact. Anthropic’s are littered with literary devices, metaphors, >>> dependent clauses, adverbs, and read like a detective novel. They are also >>> very verbose – unsurprising given they bill by the token. >>> >>> I’d like to propose a norm for how we approach comments and >>> documentation in the codebase. The proposal is that all comments and >>> documentation should maintain our flat and neutral tone, and read as >>> indistinguishable from human committer authorship. I’d like for us to >>> normalize watching for this in review as well to maintain the quality of >>> our in-tree documentation. >>> >>> I’d also like to propose removing the Claudish in CASSANDRA-21462 and >>> replacing it with comments and documentation that are in step with how we >>> write. >>> >>> Interested in others’ thoughts on this. >>> >>> – Scott >>> >>> [ This thread’s topic is limited to literary style in comments and >>> documentation. If there are other topics related to model authorship or the >>> patch above, please discuss them on a separate thread. ] >>> >>> >>> >>>
