I love this conversation.

Below is what I think and I talk in general, I haven’t looked at the
particular PR mentioned.

I think adding a skill around writing docs/comments and actually settling
on some directions about what we aim for as a community, so new people can
also learn quickly, not just by vibes, reading around - sounds great.
Thanks David, for bringing it up.

“My 2 cents would be, when we review comments/documentation, the questions
we ask ourselves should be:
1. Is it correct?
2. Is it clear?
3. Is it concise?
And any other good qualities we have been looking for in documentation
since pre-AI time.

But not “does it look like AI-generated?” If the AI-generated doc is in
good quality, what’s the problem?”

I agree on this. I catch myself cutting/editing a lot of AI-generated
docs/comments for good reasons stated here. But also, as a non-native
speaker, I love how AI can quickly improve my writing, my grammar.
And I wanted to mention - sometimes I ask it to improve what I wrote, then
it feels a bit different style than the one I would use, but it is still
accurate and makes sense - then I leave it as is but I also stress even
more for others that AI helped, especially with the writing, but I verified
it and I fixed it. It is different but verified and good, so I’d say it
serves the needs of documenting more, which Chris mentioned we were missing
a lot in time.
There is also the psychological element - we all don’t want to allow slop,
so the moment we see something that doesn’t sound exactly as the person who
submitted it, many times a red lamp turns on, and we catch ourselves
wondering how much the person who submitted it looked into it etc., but
that doesn’t mean the submission is immediately bad.

*This email was fully Ekaterina-written. :-) (I am sure there are some
grammar mistakes to prove it)*

On Sun, 23 Aug 2026 at 15:17, Jane H <[email protected]> wrote:

> Thank you for bringing this up!
>
> My 2 cents would be, when we review comments/documentation, the questions
> we ask ourselves should be:
> 1. Is it correct?
> 2. Is it clear?
> 3. Is it concise?
> And any other good qualities we have been looking for in documentation
> since pre-AI time.
>
> But not “does it look like AI-generated?” If the AI-generated doc is in
> good quality, what’s the problem?
>
> So, I agree that we should use the word API instead of “the load bearing
> seam”, cuz it’s not clear. I also agree that we should get rid of the
> verbose comments, cuz they are not concise. But I don’t think we need to
> bother to make it looks human-written, like getting rid of the dash - as
> long as it’s still correct, clear, and concise.
>
> Jane
>
>
>
>
> On Sun, Aug 23, 2026 at 11:45 Chris Lohfink <[email protected]> wrote:
>
>> I guess I have a bit of the opposite view. After years of no comments
>> and vague documentation we have correct albeit verbose PRs coming up. While
>> I think it's perfectly acceptable to create PRs or make recommendations
>> that do things like make comments more readable, I am -1 to the idea of
>> restricting or having requirements verbiage to being a nebulous "better". I
>> think that also leads to a path that can discriminate against non native
>> speakers for difficult to understand English in their commits.
>>
>> Chris
>>
>> On Sun, Aug 23, 2026 at 1:35 PM David Capwell <[email protected]> wrote:
>>
>>> Maybe a native speaker can quickly digest and filter out all witticisms,
>>> but for folks like me it might be challenging.
>>>
>>>
>>> We can’t… there is a reason that 2 skill patterns have taken hold: speak
>>> to me like i’m 5, and asd-ste100… in short it’s word vomit that no one
>>> understands.  What’s worse is that trying to understand it confuses humans
>>> and Claude; leading to worse outputs than if the comments were just thrown
>>> away.
>>>
>>>
>>> Sent from my iPhone
>>>
>>> On Aug 23, 2026, at 6:23 AM, Alex Petrov <[email protected]> wrote:
>>>
>>> 
>>> Thank you for bringing the subject of LLM-generated prose. I'd like to
>>> first mention that I only had a cursory glance at 21452/21462, and my
>>> comment is about LLM-generated prose in general rather than this patch in
>>> particular.
>>>
>>> My recent impression after reading an LLM-generated README in an OSS
>>> project (albeit not in Cassandra) was the feeling that I have started
>>> reading mid-paragraph, and the writing itself was witty and confusing,
>>> which ultimately lead me to closing the browser window rather than
>>> bothering to digest that text.
>>>
>>> I tend to remove all LLM-generated comments from my own code. For me,
>>> clarity goes beyond comments in code: any artefact that another human is
>>> going to read (for example, a code-review post, or bug
>>> analysis/investigation), needs to pass a readability bar. Having that said,
>>> I hate to admit that my own writing style might have also changed for worse
>>> over the course of this year.
>>>
>>> An additional reason (besides consistency with our existing comments and
>>> documentation) for having simpler language in documentation and code is the
>>> fact that for many of us English is not a first (and for some, not even a
>>> second) language. Maybe a native speaker can quickly digest and filter out
>>> all witticisms, but for folks like me it might be challenging.
>>>
>>> +1 for simplifying the language for both AI- and human generated prose.
>>>
>>>
>>> On Sun, Aug 23, 2026, at 2:01 AM, [email protected] wrote:
>>>
>>> Hi all,
>>>
>>> Anthropic’s current models are famous for generating overwrought
>>> metaphorical constructions like “the load-bearing seam” when referring to
>>> something as simple as an interface. r/ClaudeAI has dubbed this manner of
>>> speaking “Claudish.” Many users (including myself) have elaborate user
>>> prompts that try to tame the model, while others go as far as passing
>>> Opus/Fable-generated output through a competitor’s model to untangle it.
>>>
>>> Like many, I find reading Claudish grating and artificial  - like the
>>> taste of a Sweet ’N Low packet (aspartame), or listening to a 48kbps MP3
>>> dominated by compression artifacts.
>>>
>>> I’d like to start a discussion about project norms regarding
>>> model-generated comments and documentation in our codebase, largely
>>> prompted by the merge of CASSANDRA-21462 (5b34068).
>>>
>>> I open with my gratitude for work to validate and harden cursor-based
>>> compaction. My local measurements land it between 1.7 - 2.4x the throughput
>>> of legacy iterator-based compaction – a stunning improvement that will make
>>> Cassandra faster and more stable. I also appreciate the focus on
>>> correctness and validation in this work, as it surfaced and resolved
>>> several serious issues.
>>>
>>> The concern it prompts for me is that the commit marks the first
>>> introduction of Claudish into the codebase, and quite a lot of it.
>>>
>>> Examples in the first 1/3 of the patch include:
>>>
>>> – The zero case is load-bearing rather than an optimisation
>>> – The seam is EVENT-shaped because {@link UnfilteredDescriptor}s are
>>> transient
>>> – The row-side analogue of cell reconciliation, is load-bearing in the
>>> cursor's row
>>> – A decoder defect is as likely as on-disk damage here, but that's the
>>> same ambiguity
>>> – The mirror has already drifted from the upstream serializer once
>>> _ The precondition is asserted rather than assumed
>>> – The cursor compaction path and the reference path reach one decision.
>>> They did not always: the cursor carried a hand-mirrored copy
>>>
>>> The linguistic style of Anthropic’s models is sharply out of step with
>>> comments in Cassandra’s codebase. Our comments are concise, flat, and
>>> matter-of-fact. Anthropic’s are littered with literary devices, metaphors,
>>> dependent clauses, adverbs, and read like a detective novel. They are also
>>> very verbose – unsurprising given they bill by the token.
>>>
>>> I’d like to propose a norm for how we approach comments and
>>> documentation in the codebase. The proposal is that all comments and
>>> documentation should maintain our flat and neutral tone, and read as
>>> indistinguishable from human committer authorship. I’d like for us to
>>> normalize watching for this in review as well to maintain the quality of
>>> our in-tree documentation.
>>>
>>> I’d also like to propose removing the Claudish in CASSANDRA-21462 and
>>> replacing it with comments and documentation that are in step with how we
>>> write.
>>>
>>> Interested in others’ thoughts on this.
>>>
>>> – Scott
>>>
>>> [ This thread’s topic is limited to literary style in comments and
>>> documentation. If there are other topics related to model authorship or the
>>> patch above, please discuss them on a separate thread. ]
>>>
>>>
>>>
>>>

Reply via email to