Hello and a good day.
Steffen Nurpmeso wrote in
<20260324144838.Lop5n115@steffen%sdaoden.eu>:
It has been over two months since i sent this message, so i want
to give a follow-up. I am not yet ready with the draft, in fact
i did start working on it not sooner but last week.
I try hard to finish it by the end of the next week.
I will not post here when it is submitted.
I will ask -IESG@ to publish it as an individual submission.
My last message of mine to this list, i have nothing to add to
that other than what is in my draft. I do not want to be
mentioned in any thinkable way in the published outcome of what
you do.
To reiterate what DKIM/ACDC was, is, and will be.
And the destructive IETF approaches besides plain SMTP are not.
- The IETF built a galaxy of thousands around the RFC 6376 saying
that one "<em>SHOULD NOT determine message acceptability based
solely on a lack of any signature or on an unverifiable
signature; such rejection would cause severe interoperability
problems</em>".
This is wrong.
DKIM/ACDC goes the road of all security protocols by relying
completely upon cryptographic signatures.
- My personal impression of the often repeated saying of a person
educated so well, and with likely excellent grading, who swam in
this world of excellency, hearing here, and there, with such
outstanding teaching professors, like Wei Chuang, namely
"aligned with DMARC from the start" (iterated), made me little
idiot from the province think of Goodhart's law:
Any observed statistical regularity will tend to collapse
once pressure is placed upon it for control purposes.
DMARC is devastating. ARC.
DKIM/ACDC goes the road of all security protocols by relying
completely upon cryptographic signatures.
(Why would anyone do anything but rejecting a message that has
been tampered with? Why??)
- DKIM/ACDC is implemented in *one* program.
*One* program has access to the private key.
This WG, as i understood from reading last, wants to either
spread access to the private key to many programs, which is
a security disaster, or they want to introduce application
specific keys, but this belongs into a AKIM working group, it is
AKIM1, and not DKIM2.
- DKIM/ACDC has no business with messing up SMTP.
We do not introduce any control flow "enforcings", let alone
bizarre ones like "donotexplode".
From some younger internet history list posting:
I recall many Usenet sites supported this feature, but don’t
remember which ones off the top of my head. With sendmail,
one would put an entry into the ‘aliases’ file such as:
news-alias: |"/usr/bin/inews -h"
Is all that machinery now subject to being destroyed by the
IETF, after all the world of aliases had been ruined before?
How shall those things be enforced? In a billion scripts?
No, this WG wants to destruct the billion scripts.
- DKIM/ACDC restricts itself as much as possible to avoid any
"information Disclosure in trace fields" (and everywhere else).
With it, a completely anonymous trace header path is exactly
that.
We do not hinder or constrain SMTP in any way.
- DKIM/ACDC is a backward compatible extension of DKIMv1.
+ It now introduces a flag day (t= value) after which backward
compatibility is lost, however.
For one because it "overcomes" the very long h= list by using
a bitset, as i have posted already, and as below;
after the flag day v1 will no longer verify since most headers
will no longer be legacy-reiterated in h=.
Second because it will loose the RSA algorithm.
The next iteration turns to use elliptic curve for any "glue
record", DKIX-DC:, DKIX-AC: thus; plus DKIX-Signature:.
Only until the flag day a single DKIM-Signature: exists, for
compatibility purposes, of whatever algorithm is chosen.
- It protects against replay attacks, and backscatter bounces.
- It creates a set of differential changes that can be applied in
reverse order, cryptographically safe, to be able to verify
elder signatures.
+ I have created a very simple, very (very) fast line-based
differential algorithm, as below. The nice property of the
BSDiff patch format, that is comprehensive security checks,
perfectly spaced buffer allocations, direct usability of the
result buffer for verification purposes, is retained.
One can choose, also dynamically, in between the algorithms.
Examples and algorithm as below.
+ Here the rationale for using the line based textual diff:
Like SMTP DKIM does not know about MIME,
it treats the body (and the header fields, in that respect) as
<tt>CRLF</tt> terminated lines of bytes,
with certain byte-based ("relaxed") normalizations applied on top.
MIME reencoding (may) happen along the message path:
line lengths, the MIME content-transfer-encoding type,
as well as character sets, and more, can theoretically be in flux.
With the advent of tracking of differential changes it is expected
that software becomes smarter, by adapting more to what exists in
messages, instead of performing "brute force modifications".
For example, a service provider that rewrites URIs within messages
can ensure that the line lengths of a base64 formatted input are
preserved after the rewrite, as base64, and with the underlaying
character set being unchanged, and that, where modifications took
place, the lengths of only the modified lines are adjusted so to
keep the differential changes as minimal as possible.
In such a world a very fast linewise algorithm is sufficient.
(But otherwise i think it should really be a BSDiff thing.)
+ Only ZLIB will remain as a compression method. Whereas
certain numbers are worse in the tests, it must be said that
MIME parts are often images and the like, and therefore highly
compressed already. Other compressors gain not really
anything in practice, but create further dependencies.
- I have seen the DKIM2 of this WG chooses "simple" instead of
"relaxed" for body normalization. I have this rationale for the
other type, but i am at odds (a bit).
The "extremely crude ASCII Art attacks" mentioned in
DKIM<xref target="RFC6376"/>
section 8.1 are considered to be a rather artificial attack vector.
Furthermore the statement of DKIM section
5.4.1. "Recommended Signature Content"
that e-commerce sites (etc) <em>will generally prefer "simple"
canonicalization</em>
does not match reality in the author's experience.
Almost exclusively it is relied on enriched format messages
with dedicated whitespace rules, like HTML,
or even more refined (compressable, encryptable, optionally
interactive, scripted) formats, like PDF.
In addition, and foremost, choosing space-preserving MIME content
transfer encodings is the, used, natural choice, as appropriate.
As a personal comment, because rarely used in practice, there exist
possibilities to sign and/or encrypt actual message content to
ensure its privacy on a per recipient base, like S/MIME and PGP,
which use dedicated, replicable normalization algorithms to protect
their content.
In operational reality the "relaxed" normalization is by far the
most commonly used form, percent-wise it possibly can be estimated
in the high nineties.
It follows that whitespace granularity that exact does not matter
for domain signatures,
and if reduction to a single algorithm is desired,
"relaxed" is the only viable form for header fields (anyway),
and using it for message bodies is a most widely accepted variant.
Yes. Who uses simple? IETF (and thus? IANA) do not count, it
does very bitter things to email. OpenBSD, yes, Viktor
Dukhovni, and then we already come to certain so-called "nerds".
I mean "go and use privacy!", use S/MIME and PGP if you want
that, but this is a domain-key, and for that relax serves its
purpose very well.
On the other hand i am at odds, but the above is not polemic.
- DKIM/ACDC will "go DKIX" completely.
This means only one "DKIM-" remains, and that is a single
signature that is compatible. Everything else uses a DKIX-
prefix, and is not compatible -- it does not need to.
- DKIM/ACDC uses only elliptic curve algorithms.
In fact it will use the new adaed25519 approach, not RFC 8463 as
such; rationale:
..
Different to plain DKIM DKIM/ACDC requires to keep the
normalized content of header fields around for consumption
by the differential changes algorithm.
The immediate consumption of all input by "hash-alg",
as was and is implemented by certain (not all) DKIM software,
can therefore no longer be a primary design goal.
..
As explained in DKIM section 3.7,
many digital signature APIs combine hashing and signature creation.
RFC 8463's "ed25519" being rarely implemented to this day could also
be based on the fact that existing code bases using these APIs require
code flow adjustments for being compatible with it.
And given that immediate consumption of header field data is no longer
possible,
and because in the future DKIM administrative data will shrink in size,
as shown below, this document makes extensive use of "adaed25519".
- DKIM/ACDC will use a bitset to express which of the header
fields of a known database of headers were present at signing
time. These are all sealed (oversigned). Only additional user
header fields remain in "h=". *However*, before the flag day,
to remain compatibility with DKIMv1, the single DKIM-Signature:
that is stored iterates all header fields also in "h=".
The bitsets stores 5 bits per bytes.
Personal mails likely require 4, list emails 6 byte bitsets.
(The C code posted last had a bug in that it missed an error
case; it now also supports case-insensitivity.)
- Adds new flags "C" (interest in collection and (periodical)
report of statistic informations, otherwise out of scope), and
"T" and "t":
the "T" and "t" flags are meant to adapt to operational reality.
There "trusted proof points" are hired to handle email, to apply
all the necessary checks for and removal of spam, malicious,
dangerous, or otherwise undesired message (MIME part) content,
before passing the results further to their real recipients.
As of today only the "equivalent of T" is a known mode of operation;
DKIM/ACDC, however, allows for a new business model via "t":
the "trusted proof point" readily prepares messages just like today,
but also creates and includes a DKIM-DC: header field to undo these
modifications, as well as keeping all elder DKIM-DC: header fields
intact.
(Read: simply through the normal DKIM/ACDC mode of operation,
except for setting the "t" flag in addition.)
Turning a "t" message into a "T" message practically means nothing
but removing the DKIM-DC: header fields:
an operation that can fastly and safely be performed by simplemost
command line utilities or scripting languages,
thanks to the plain-text nature of
SMTP<xref target="RFC5321"/>,
of
IMF<xref target="RFC5322"/>
messages.
..
<em>Informative remark:</em>
because DKIM-DC: header fields are covered by the "ddch=" hash,
removing them still allows for successful DKIM-Signature:
verification, simply by trusting the original "ddch=" checksum.
DKIM/ACDC's "t" flag allows customers to perform a complete "R"
reputation check on data delivered by "trusted proof points".
It is only their users verifying "I" ingress signatures who have
no option but putting trust into "ddch=" hashes.
..
<em>Example:</em>
In the (real-world) scenario of direct T(ransport) L(ayer) S(ecurity)
protected connections from the customer to the "trusted proof point",
and only controlled (authenticated) storage and message access on
customer host(s),
end users would still only be able to see the proofed message
variant, without being able to apply differential changes.
However, (to be written or extended) message access software could
be allowed to perform certain operations on the "original" message
content, for example to restore certain allowed header fields.
It could, for example, restore an original From: header field as
created by the message sender, but which was rewritten along the
message path, for example, by mailing-lists.
(To be remarked that, also alternatively, the
Author Header Field<xref target="RFC9057"/>
can play a crucial role in this,
whether in a restored, or created, or "generally provided",
variant)
- The DNS entry will be a bit different than before. No more
limits etc, this was a maintenance disaster.
And quite a bit. But i think the email is long enough now.
I say Ciao! already here,
..may at least the force not be with you, so to say..
Patria o Muerte.
Many greetings.
Appendix I. Header Field Database:
By using a bitset for these headers, storing five (5) bits per
byte, a normal personal email requires four (4) bytes, a list
email requires six (6) bytes for the bitset that replaces "h=",
and covers all instances of the announced header fields, plus "a
seal".
=HFDB
=1
# v1 sorted for normal 4, list 6
# RFC 5322, 3.6 / 5322-bis, I.
date 1
from 2
sender 3
reply-to 4
to 5
cc 6
bcc 7
message-id 8
in-reply-to 9
references 10
subject 11
# RFC 9057
author 12
# RFC 2045
mime-version 13
content-type 14
content-transfer-encoding 15
# (rest usually not in main header)
content-id 16
content-description 17
# RFC 1806
content-disposition 18
# draft-ietf-drums-mail-followup-to-00
mail-followup-to 19
# draft-josefsson-openpgp-mailnews-header-07
openpgp 20
# RFC 2369
list-id 21
list-help 22
list-subscribe 23
list-unsubscribe 24
list-post 25
list-owner 26
list-archive 27
# RFC 5064
archived-at 28
# RFC 5322, 3.6 / 5322-bis, II.
resent-from 29
resent-sender 30
resent-to 31
resent-cc 32
resent-date 33
resent-message-id 34
resent-bcc 35
comments 36
keywords 37
=/1
=/HFDB
Appendix II. Line based difference algorithm
Examples; times on this box vary a bit because of the overall
load :), these are neither the fastest nor the slowest ones,
but some numbers "picked from the runs". I compare against
the 70s ed(1) script output of diff(1), because it seems to
reflect somewhat the algorithm what this WG develops. Tests
suggest that even base64 encoded the ed(1) script is larger,
without having any security or convenience property.
# Large roff manual (428420 vs rewritten 390770 bytes)
$ diff -e .B1 .B2 | wc -c
241233
# BSDiff: data: ctrl=54816 (4568 entries) diff=304307 extra=124113
# 74830 result bytes; Code 0:198 secs, ZLIB I/O 0:122 secs
# (BZ2 67368 / 0:022, XZ 65216 / 0:137, ZSTD 69444 / 0:144)
# Textual: data: ctrl=54456 (4538 entries) diff=202309 extra=226111
# 97355 result bytes; Code 0:004 secs, ZLIB I/O 0:098 secs
# (BZ2 81943 / 0:053, XZ 82364 / 0:253, ZSTD 86767 / 0:213)
# Email, sent to/received from ML (4621 vs 5309 bytes):
$ diff -e .S1 .S2 | wc -c
2326
# BSDiff: data: ctrl=216 (18 entries) diff=3554 extra=1067
# 857 result bytes; Code 0:001 secs, ZLIB I/O 0:000 secs
# (BZ2 950 / 0:001, XZ 884 / 0:126, ZSTD 875 / 0:107)
# Textual: data: ctrl=108 (9 entries) diff=3043 extra=1578
# 1089 result bytes; Code 0:000 secs, ZLIB I/O 0:000 secs
# (BZ2 1260 / 0:001, XZ 1144 / 0:127, ZSTD 1116 / 0:126)
Algorithm.
<t>
Initialize a "list of lines" that can be iterated over
uni-directionally to 0.
While there is still input data from the target (egress) data set:
</t><ol><li>
Search for <tt>LF</tt> line feed.
If found, advance over it,
otherwise use the entire remaining input.
</li><li>
Verify the resulting data length fits into 31-bit, minus 1.
</li><li>
Create a hash of the data that is proof against complexity attacks.
</li><li>
Put the data at the end of the "list of lines".
</li></ol><t>
Initialize a "hashmap" to 0.
If any, store all members of the "list of lines" therein,
indexed by their hashes.
</t><t>
Initialize an "absolute position",
a "current length of differences",
and a "current length of extra data" to 0.
Initialize a "former line" to 0.
While there is still input data from the source (ingress) data set:
</t><ol><li>
Search for <tt>LF</tt> line feed.
If found, advance over it,
otherwise use the entire remaining input.
</li><li>
Verify the resulting data length fits into 31-bit, minus 1.
</li><li>
If the "hashmap" is not 0:
<ol><li>
Create a hash of the data (with the same algorithm as above).
</li><li>
If the "former line" is not 0,
check whether its next line (in the "list of lines"), if any,
matches the current data.
If so,
make the next line the "former line",
verify the resulting "overall length of differences"
will fit in 31-bit, minus 1;
add that many bytes to the "absolute position",
add that many bytes to the "current length of differences",
add that many zero bytes to "differences".
</li><li>
</li><li>
Otherwise search "hashmap" for the hash.
If a match is found,
make it the "former line";
then if either the "current length of differences" is not 0,
or the "current length of extra data" is not 0,
or if no control tuple has yet been dumped,
then dump a control tuple first,
and set the "current length of extra data" to 0;
if the offset of the start of the "former line"
(within the target (egress) data set)
does not equal the "absolute position",
ensure the relative seek of the formerly dumped control tuple
is set to the subtraction of the start of the "former line"
and the "absolute position";
set the "current length of differences" to that many bytes,
set the "absolute position" to the start of the "former line"
plus that many bytes,
add that many zero bytes to "differences".
(Optimization: dumping an initial control tuple is not necessary
if no relative seek is to be applied.)
</li></ol>
</li><li>
Otherwise the line is extra data.
verify the resulting "overall length of extra data"
will fit in 31-bit, minus 1;
add that many bytes to the "current length of extra data",
copy that many bytes to "extra data".
</li></ol><t>
After all the data has been worked,
then if either the "current length of differences" is not 0,
or the "current length of extra data" is not 0,
then dump a control tuple.
</t><t>
Ensure the relative seek of the last control tuple, if any, is 0.
</t>
--steffen
|
|Der Kragenbaer, The moon bear,
|der holt sich munter he cheerfully and one by one
|einen nach dem anderen runter wa.ks himself off
|(By Robert Gernhardt)
_______________________________________________
Ietf-dkim mailing list -- [email protected]
To unsubscribe send an email to [email protected]