Hello and a good day.

Steffen Nurpmeso wrote in
 <20260324144838.Lop5n115@steffen%sdaoden.eu>:

It has been over two months since i sent this message, so i want
to give a follow-up.  I am not yet ready with the draft, in fact
i did start working on it not sooner but last week.
I try hard to finish it by the end of the next week.
I will not post here when it is submitted.
I will ask -IESG@ to publish it as an individual submission.
My last message of mine to this list, i have nothing to add to
that other than what is in my draft.  I do not want to be
mentioned in any thinkable way in the published outcome of what
you do.

To reiterate what DKIM/ACDC was, is, and will be.
And the destructive IETF approaches besides plain SMTP are not.

- The IETF built a galaxy of thousands around the RFC 6376 saying
  that one "<em>SHOULD NOT determine message acceptability based
  solely on a lack of any signature or on an unverifiable
  signature; such rejection would cause severe interoperability
  problems</em>".

  This is wrong.
  DKIM/ACDC goes the road of all security protocols by relying
  completely upon cryptographic signatures.

- My personal impression of the often repeated saying of a person
  educated so well, and with likely excellent grading, who swam in
  this world of excellency, hearing here, and there, with such
  outstanding teaching professors, like Wei Chuang, namely
  "aligned with DMARC from the start" (iterated), made me little
  idiot from the province think of Goodhart's law:
    Any observed statistical regularity will tend to collapse
    once pressure is placed upon it for control purposes.

  DMARC is devastating.  ARC.
  DKIM/ACDC goes the road of all security protocols by relying
  completely upon cryptographic signatures.
  (Why would anyone do anything but rejecting a message that has
  been tampered with?  Why??)

- DKIM/ACDC is implemented in *one* program.
  *One* program has access to the private key.

  This WG, as i understood from reading last, wants to either
  spread access to the private key to many programs, which is
  a security disaster, or they want to introduce application
  specific keys, but this belongs into a AKIM working group, it is
  AKIM1, and not DKIM2.

- DKIM/ACDC has no business with messing up SMTP.
  We do not introduce any control flow "enforcings", let alone
  bizarre ones like "donotexplode".
  From some younger internet history list posting:

    I recall many Usenet sites supported this feature, but don’t
    remember which ones off the top of my head.  With sendmail,
    one would put an entry into the ‘aliases’ file such as:

      news-alias: |"/usr/bin/inews -h"

  Is all that machinery now subject to being destroyed by the
  IETF, after all the world of aliases had been ruined before?
  How shall those things be enforced?  In a billion scripts?
  No, this WG wants to destruct the billion scripts.

- DKIM/ACDC restricts itself as much as possible to avoid any
  "information Disclosure in trace fields" (and everywhere else).
  With it, a completely anonymous trace header path is exactly
  that.

  We do not hinder or constrain SMTP in any way.

- DKIM/ACDC is a backward compatible extension of DKIMv1.

  + It now introduces a flag day (t= value) after which backward
    compatibility is lost, however.

    For one because it "overcomes" the very long h= list by using
    a bitset, as i have posted already, and as below;
    after the flag day v1 will no longer verify since most headers
    will no longer be legacy-reiterated in h=.

    Second because it will loose the RSA algorithm.
    The next iteration turns to use elliptic curve for any "glue
    record", DKIX-DC:, DKIX-AC: thus; plus DKIX-Signature:.
    Only until the flag day a single DKIM-Signature: exists, for
    compatibility purposes, of whatever algorithm is chosen.

- It protects against replay attacks, and backscatter bounces.

- It creates a set of differential changes that can be applied in
  reverse order, cryptographically safe, to be able to verify
  elder signatures.

  + I have created a very simple, very (very) fast line-based
    differential algorithm, as below.  The nice property of the
    BSDiff patch format, that is comprehensive security checks,
    perfectly spaced buffer allocations, direct usability of the
    result buffer for verification purposes, is retained.
    One can choose, also dynamically, in between the algorithms.
    Examples and algorithm as below.

  + Here the rationale for using the line based textual diff:

      Like SMTP DKIM does not know about MIME,
      it treats the body (and the header fields, in that respect) as
      <tt>CRLF</tt> terminated lines of bytes,
      with certain byte-based ("relaxed") normalizations applied on top.

      MIME reencoding (may) happen along the message path:
      line lengths, the MIME content-transfer-encoding type,
      as well as character sets, and more, can theoretically be in flux.

      With the advent of tracking of differential changes it is expected
      that software becomes smarter, by adapting more to what exists in
      messages, instead of performing "brute force modifications".

      For example, a service provider that rewrites URIs within messages
      can ensure that the line lengths of a base64 formatted input are
      preserved after the rewrite, as base64, and with the underlaying
      character set being unchanged, and that, where modifications took
      place, the lengths of only the modified lines are adjusted so to
      keep the differential changes as minimal as possible.

      In such a world a very fast linewise algorithm is sufficient.

    (But otherwise i think it should really be a BSDiff thing.)

  + Only ZLIB will remain as a compression method.  Whereas
    certain numbers are worse in the tests, it must be said that
    MIME parts are often images and the like, and therefore highly
    compressed already.  Other compressors gain not really
    anything in practice, but create further dependencies.

- I have seen the DKIM2 of this WG chooses "simple" instead of
  "relaxed" for body normalization.  I have this rationale for the
  other type, but i am at odds (a bit).

    The "extremely crude ASCII Art attacks" mentioned in
    DKIM<xref target="RFC6376"/>
    section 8.1 are considered to be a rather artificial attack vector.

    Furthermore the statement of DKIM section
    5.4.1. "Recommended Signature Content"
    that e-commerce sites (etc) <em>will generally prefer "simple"
    canonicalization</em>
    does not match reality in the author's experience.

    Almost exclusively it is relied on enriched format messages
    with dedicated whitespace rules, like HTML,
    or even more refined (compressable, encryptable, optionally
    interactive, scripted) formats, like PDF.

    In addition, and foremost, choosing space-preserving MIME content
    transfer encodings is the, used, natural choice, as appropriate.

    As a personal comment, because rarely used in practice, there exist
    possibilities to sign and/or encrypt actual message content to
    ensure its privacy on a per recipient base, like S/MIME and PGP,
    which use dedicated, replicable normalization algorithms to protect
    their content.

    In operational reality the "relaxed" normalization is by far the
    most commonly used form, percent-wise it possibly can be estimated
    in the high nineties.
    It follows that whitespace granularity that exact does not matter
    for domain signatures,
    and if reduction to a single algorithm is desired,
    "relaxed" is the only viable form for header fields (anyway),
    and using it for message bodies is a most widely accepted variant.

  Yes.  Who uses simple?  IETF (and thus? IANA) do not count, it
  does very bitter things to email.  OpenBSD, yes, Viktor
  Dukhovni, and then we already come to certain so-called "nerds".
  I mean "go and use privacy!", use S/MIME and PGP if you want
  that, but this is a domain-key, and for that relax serves its
  purpose very well.
  On the other hand i am at odds, but the above is not polemic.

- DKIM/ACDC will "go DKIX" completely.
  This means only one "DKIM-" remains, and that is a single
  signature that is compatible.  Everything else uses a DKIX-
  prefix, and is not compatible -- it does not need to.

- DKIM/ACDC uses only elliptic curve algorithms.
  In fact it will use the new adaed25519 approach, not RFC 8463 as
  such; rationale:

  ..
    Different to plain DKIM DKIM/ACDC requires to keep the
    normalized content of header fields around for consumption
    by the differential changes algorithm.

    The immediate consumption of all input by "hash-alg",
    as was and is implemented by certain (not all) DKIM software,
    can therefore no longer be a primary design goal.
  ..
    As explained in DKIM section 3.7,
    many digital signature APIs combine hashing and signature creation.

    RFC 8463's "ed25519" being rarely implemented to this day could also
    be based on the fact that existing code bases using these APIs require
    code flow adjustments for being compatible with it.

    And given that immediate consumption of header field data is no longer
    possible,
    and because in the future DKIM administrative data will shrink in size,
    as shown below, this document makes extensive use of "adaed25519".

- DKIM/ACDC will use a bitset to express which of the header
  fields of a known database of headers were present at signing
  time.  These are all sealed (oversigned).  Only additional user
  header fields remain in "h=".  *However*, before the flag day,
  to remain compatibility with DKIMv1, the single DKIM-Signature:
  that is stored iterates all header fields also in "h=".

  The bitsets stores 5 bits per bytes.
  Personal mails likely require 4, list emails 6 byte bitsets.
  (The C code posted last had a bug in that it missed an error
  case; it now also supports case-insensitivity.)

- Adds new flags "C" (interest in collection and (periodical)
  report of statistic informations, otherwise out of scope), and
  "T" and "t":

    the "T" and "t" flags are meant to adapt to operational reality.
    There "trusted proof points" are hired to handle email, to apply
    all the necessary checks for and removal of spam, malicious,
    dangerous, or otherwise undesired message (MIME part) content,
    before passing the results further to their real recipients.

    As of today only the "equivalent of T" is a known mode of operation;
    DKIM/ACDC, however, allows for a new business model via "t":
    the "trusted proof point" readily prepares messages just like today,
    but also creates and includes a DKIM-DC: header field to undo these
    modifications, as well as keeping all elder DKIM-DC: header fields
    intact.
    (Read: simply through the normal DKIM/ACDC mode of operation,
    except for setting the "t" flag in addition.)

    Turning a "t" message into a "T" message practically means nothing
    but removing the DKIM-DC: header fields:
    an operation that can fastly and safely be performed by simplemost
    command line utilities or scripting languages,
    thanks to the plain-text nature of
    SMTP<xref target="RFC5321"/>,
    of
    IMF<xref target="RFC5322"/>
    messages.
  ..
    <em>Informative remark:</em>
    because DKIM-DC: header fields are covered by the "ddch=" hash,
    removing them still allows for successful DKIM-Signature:
    verification, simply by trusting the original "ddch=" checksum.

    DKIM/ACDC's "t" flag allows customers to perform a complete "R"
    reputation check on data delivered by "trusted proof points".
    It is only their users verifying "I" ingress signatures who have
    no option but putting trust into "ddch=" hashes.
  ..
    <em>Example:</em>
    In the (real-world) scenario of direct T(ransport) L(ayer) S(ecurity)
    protected connections from the customer to the "trusted proof point",
    and only controlled (authenticated) storage and message access on
    customer host(s),
    end users would still only be able to see the proofed message
    variant, without being able to apply differential changes.

    However, (to be written or extended) message access software could
    be allowed to perform certain operations on the "original" message
    content, for example to restore certain allowed header fields.
    It could, for example, restore an original From: header field as
    created by the message sender, but which was rewritten along the
    message path, for example, by mailing-lists.

    (To be remarked that, also alternatively, the
    Author Header Field<xref target="RFC9057"/>
    can play a crucial role in this,
    whether in a restored, or created, or "generally provided",
    variant)

- The DNS entry will be a bit different than before.  No more
  limits etc, this was a maintenance disaster. 

And quite a bit.  But i think the email is long enough now.


I say Ciao! already here,
..may at least the force not be with you, so to say..
Patria o Muerte.

Many greetings.


Appendix I.  Header Field Database:

By using a bitset for these headers, storing five (5) bits per
byte, a normal personal email requires four (4) bytes, a list
email requires six (6) bytes for the bitset that replaces "h=",
and covers all instances of the announced header fields, plus "a
seal".

  =HFDB
  =1
  # v1 sorted for normal 4, list 6
  # RFC 5322, 3.6 / 5322-bis, I.
  date 1
  from 2
  sender 3
  reply-to 4
  to 5
  cc 6
  bcc 7
  message-id 8
  in-reply-to 9
  references 10
  subject 11
  # RFC 9057
  author 12
  # RFC 2045
  mime-version 13
  content-type 14
  content-transfer-encoding 15
  # (rest usually not in main header)
  content-id 16
  content-description 17
  # RFC 1806
  content-disposition 18
  # draft-ietf-drums-mail-followup-to-00
  mail-followup-to 19
  # draft-josefsson-openpgp-mailnews-header-07
  openpgp 20
  # RFC 2369
  list-id 21
  list-help 22
  list-subscribe 23
  list-unsubscribe 24
  list-post 25
  list-owner 26
  list-archive 27
  # RFC 5064
  archived-at 28
  # RFC 5322, 3.6 / 5322-bis, II.
  resent-from 29
  resent-sender 30
  resent-to 31
  resent-cc 32
  resent-date 33
  resent-message-id 34
  resent-bcc 35
  comments 36
  keywords 37
  =/1
  =/HFDB

Appendix II.  Line based difference algorithm

Examples; times on this box vary a bit because of the overall
load :), these are neither the fastest nor the slowest ones,
but some numbers "picked from the runs".  I compare against
the 70s ed(1) script output of diff(1), because it seems to
reflect somewhat the algorithm what this WG develops.  Tests
suggest that even base64 encoded the ed(1) script is larger,
without having any security or convenience property.

  # Large roff manual (428420 vs rewritten 390770 bytes)
  $ diff -e .B1 .B2 | wc -c
  241233
  #  BSDiff: data: ctrl=54816 (4568 entries) diff=304307 extra=124113
  #          74830 result bytes; Code 0:198 secs, ZLIB I/O 0:122 secs
  #          (BZ2 67368 / 0:022, XZ 65216 / 0:137, ZSTD 69444 / 0:144)
  # Textual: data: ctrl=54456 (4538 entries) diff=202309 extra=226111
  #          97355 result bytes; Code 0:004 secs, ZLIB I/O 0:098 secs
  #          (BZ2 81943 / 0:053, XZ 82364 / 0:253, ZSTD 86767 / 0:213)
  
  # Email, sent to/received from ML (4621 vs 5309 bytes):
  $ diff -e .S1 .S2 | wc -c
  2326
  #  BSDiff: data: ctrl=216 (18 entries) diff=3554 extra=1067
  #          857 result bytes; Code 0:001 secs, ZLIB I/O 0:000 secs
  #          (BZ2 950 / 0:001, XZ 884 / 0:126, ZSTD 875 / 0:107)
  # Textual: data: ctrl=108 (9 entries) diff=3043 extra=1578
  #          1089 result bytes; Code 0:000 secs, ZLIB I/O 0:000 secs
  #          (BZ2 1260 / 0:001, XZ 1144 / 0:127, ZSTD 1116 / 0:126)

Algorithm.

  <t>
    Initialize a "list of lines" that can be iterated over
    uni-directionally to 0.
    While there is still input data from the target (egress) data set:
  </t><ol><li>
    Search for <tt>LF</tt> line feed.
    If found, advance over it,
    otherwise use the entire remaining input.
  </li><li>
    Verify the resulting data length fits into 31-bit, minus 1.
  </li><li>
    Create a hash of the data that is proof against complexity attacks.
  </li><li>
    Put the data at the end of the "list of lines".
  </li></ol><t>
    Initialize a "hashmap" to 0.
    If any, store all members of the "list of lines" therein,
    indexed by their hashes.
  </t><t>
    Initialize an "absolute position",
    a "current length of differences",
    and a "current length of extra data" to 0.
    Initialize a "former line" to 0.
    While there is still input data from the source (ingress) data set:
  </t><ol><li>
    Search for <tt>LF</tt> line feed.
    If found, advance over it,
    otherwise use the entire remaining input.
  </li><li>
    Verify the resulting data length fits into 31-bit, minus 1.
  </li><li>
    If the "hashmap" is not 0:
    <ol><li>
      Create a hash of the data (with the same algorithm as above).
    </li><li>
      If the "former line" is not 0,
      check whether its next line (in the "list of lines"), if any,
      matches the current data.

      If so,
      make the next line the "former line",
      verify the resulting "overall length of differences"
      will fit in 31-bit, minus 1;
      add that many bytes to the "absolute position",
      add that many bytes to the "current length of differences",
      add that many zero bytes to "differences".
    </li><li>

    </li><li>
      Otherwise search "hashmap" for the hash.

      If a match is found,
      make it the "former line";
      then if either the "current length of differences" is not 0,
      or the "current length of extra data" is not 0,
      or if no control tuple has yet been dumped,
      then dump a control tuple first,
      and set the "current length of extra data" to 0;
      if the offset of the start of the "former line"
      (within the target (egress) data set)
      does not equal the "absolute position",
      ensure the relative seek of the formerly dumped control tuple
      is set to the subtraction of the start of the "former line"
      and the "absolute position";
      set the "current length of differences" to that many bytes,
      set the "absolute position" to the start of the "former line"
      plus that many bytes,
      add that many zero bytes to "differences".

      (Optimization: dumping an initial control tuple is not necessary
      if no relative seek is to be applied.)
    </li></ol>
  </li><li>
    Otherwise the line is extra data.

    verify the resulting "overall length of extra data"
    will fit in 31-bit, minus 1;
    add that many bytes to the "current length of extra data",
    copy that many bytes to "extra data".
  </li></ol><t>
    After all the data has been worked,
    then if either the "current length of differences" is not 0,
    or the "current length of extra data" is not 0,
    then dump a control tuple.
  </t><t>
    Ensure the relative seek of the last control tuple, if any, is 0.
  </t>

--steffen
|
|Der Kragenbaer,                The moon bear,
|der holt sich munter           he cheerfully and one by one
|einen nach dem anderen runter  wa.ks himself off
|(By Robert Gernhardt)

_______________________________________________
Ietf-dkim mailing list -- [email protected]
To unsubscribe send an email to [email protected]

Reply via email to