On Fri, Aug 14, 2026 at 6:08 PM Masahiko Sawada <[email protected]> wrote:
>
> On Fri, Aug 14, 2026 at 4:33 PM Masahiko Sawada <[email protected]> wrote:
> >
> > On Thu, Aug 6, 2026 at 8:07 PM Chao Li <[email protected]> wrote:
> > >
> > >
> > >
> > > > On Aug 7, 2026, at 00:09, Masahiko Sawada <[email protected]> wrote:
> > > >
> > > > On Wed, Aug 5, 2026 at 8:40 PM Chao Li <[email protected]> wrote:
> > > >>
> > > >>
> > > >>
> > > >>> On Aug 6, 2026, at 08:19, Masahiko Sawada <[email protected]> 
> > > >>> wrote:
> > > >>>
> > > >>> On Tue, Jun 30, 2026 at 11:03 AM Haibo Yan <[email protected]> 
> > > >>> wrote:
> > > >>>>
> > > >>>> On Tue, Jun 30, 2026 at 10:53 AM Masahiko Sawada 
> > > >>>> <[email protected]> wrote:
> > > >>>>>
> > > >>>>> On Mon, Jun 29, 2026 at 5:53 PM Haibo Yan <[email protected]> 
> > > >>>>> wrote:
> > > >>>>>>
> > > >>>>>> On Mon, Jun 29, 2026 at 2:55 PM Masahiko Sawada 
> > > >>>>>> <[email protected]> wrote:
> > > >>>>>>>
> > > >>>>>>> On Sun, Jun 28, 2026 at 7:20 PM Haibo Yan <[email protected]> 
> > > >>>>>>> wrote:
> > > >>>>>>>>
> > > >>>>>>>> On Thu, Jun 25, 2026 at 3:16 PM Masahiko Sawada 
> > > >>>>>>>> <[email protected]> wrote:
> > > >>>>>>>>>
> > > >>>>>>>>> On Thu, Jun 25, 2026 at 2:31 PM Haibo Yan 
> > > >>>>>>>>> <[email protected]> wrote:
> > > >>>>>>>>>>
> > > >>>>>>>>>>
> > > >>>>>>>>>>
> > > >>>>>>>>>> On Thu, Jun 25, 2026 at 11:28 AM Masahiko Sawada 
> > > >>>>>>>>>> <[email protected]> wrote:
> > > >>>>>>>>>>>
> > > >>>>>>>>>>> Hi all,
> > > >>>>>>>>>>>
> > > >>>>>>>>>>> I'd like to propose the $subject.
> > > >>>>>>>>>>>
> > > >>>>>>>>>>> Since commit ec8719ccbfcd made hex_decode_safe() SIMD-aware, 
> > > >>>>>>>>>>> decoding
> > > >>>>>>>>>>> a run of hex digits is now fast. The attached patch reuses
> > > >>>>>>>>>>> hex_decode_safe() in the UUID input function to speed up 
> > > >>>>>>>>>>> parsing.
> > > >>>>>>>>>>>
> > > >>>>>>>>>>> We accept several textual forms of a UUID[1]. The fast path 
> > > >>>>>>>>>>> handles
> > > >>>>>>>>>>> the common ones: 32 hex digits, the canonical 8x-4x-4x-4x-12x 
> > > >>>>>>>>>>> form
> > > >>>>>>>>>>> (where "nx" means n hex digits), and either of those wrapped 
> > > >>>>>>>>>>> in
> > > >>>>>>>>>>> braces. Otherwise, it falls back to the ordinary scalar UUID 
> > > >>>>>>>>>>> parse.
> > > >>>>>>>>>>>
> > > >>>>>>>>>>> I've benchmarked the parse speed using the following query:
> > > >>>>>>>>>>>
> > > >>>>>>>>>>> CREATE TEMP TABLE u AS SELECT gen_random_uuid()::text AS t 
> > > >>>>>>>>>>> FROM
> > > >>>>>>>>>>> generate_series(1, 1000000);
> > > >>>>>>>>>>> EXPLAIN (ANALYZE, TIMING OFF) SELECT t::uuid FROM u;
> > > >>>>>>>>>>>
> > > >>>>>>>>>>> I compared the execution time of the second query, which 
> > > >>>>>>>>>>> measures
> > > >>>>>>>>>>> uuid_in() alone, with/without SIMD optimization. Here are 
> > > >>>>>>>>>>> results (the
> > > >>>>>>>>>>> median of 5 runs):
> > > >>>>>>>>>>>
> > > >>>>>>>>>>> HEAD: 208.879 ms
> > > >>>>>>>>>>> Patched: 40.983 ms
> > > >>>>>>>>>>>
> > > >>>>>>>>>>> The improvements look promising to me. But in a realistic 
> > > >>>>>>>>>>> pipeline the
> > > >>>>>>>>>>> parse is a small fraction of the work, so end-to-end gains 
> > > >>>>>>>>>>> could be
> > > >>>>>>>>>>> much smaller.
> > > >>>>>>>>>>>
> > > >>>>>>>>>>> Feedback is very welcome.
> > > >>>>>>>>>>>
> > > >>>>>>>>>> I may be missing something, but I wonder whether the fast path 
> > > >>>>>>>>>> is relying on
> > > >>>>>>>>>> slightly different input semantics from the existing UUID 
> > > >>>>>>>>>> parser.
> > > >>>>>>>>>>
> > > >>>>>>>>>> In particular, hex_decode_safe() is not a strict “32 hex 
> > > >>>>>>>>>> characters only”
> > > >>>>>>>>>> decoder.  It skips whitespace, which is fine for its existing 
> > > >>>>>>>>>> callers, but I
> > > >>>>>>>>>> don’t think UUID input should treat whitespace inside the UUID 
> > > >>>>>>>>>> body as
> > > >>>>>>>>>> ignorable.
> > > >>>>>>>>>
> > > >>>>>>>>> Good catch! hex_decode_safe() skips whitespaces so the patch 
> > > >>>>>>>>> accepts
> > > >>>>>>>>> the following UUID value, which is bad:
> > > >>>>>>>>>
> > > >>>>>>>>> select '019f00b5-7f8a-722f-b707-59f0ed25cd  '::uuid;
> > > >>>>>>>>>                uuid
> > > >>>>>>>>> --------------------------------------
> > > >>>>>>>>> 019f00b5-7f8a-722f-b707-59f0ed25cd00
> > > >>>>>>>>> (1 row)
> > > >>>>>>>>>
> > > >>>>>>>>>> Also, since hex_decode_safe() returns void, the UUID fast path
> > > >>>>>>>>>> cannot verify that exactly UUID_LEN bytes were produced.
> > > >>>>>>>>>
> > > >>>>>>>>> IIUC hex_decode_safe() does return the output length in bytes. 
> > > >>>>>>>>> So I
> > > >>>>>>>>> think we can fallback to the scalar UUID parser if
> > > >>>>>>>>> esctx.error_occurred is true or if the returned value is not 16.
> > > >>>>>>>>>
> > > >>>>>>>>
> > > >>>>>>>> You’re right, I misread that part.  Checking both 
> > > >>>>>>>> esctx.error_occurred and
> > > >>>>>>>> the returned length sounds good to me.
> > > >>>>>>>>
> > > >>>>>>>>>>
> > > >>>>>>>>>> So I think it would be safer either to pre-validate that the 
> > > >>>>>>>>>> 32 source
> > > >>>>>>>>>> characters are all hex digits before calling 
> > > >>>>>>>>>> hex_decode_safe(), or to use a
> > > >>>>>>>>>> UUID-specific strict hex decoder for this path.  After that, a 
> > > >>>>>>>>>> comment
> > > >>>>>>>>>> explaining why hex_decode_safe() is safe here would make the 
> > > >>>>>>>>>> invariant much
> > > >>>>>>>>>> clearer.
> > > >>>>>>>>>
> > > >>>>>>>>> IIUC hex_decode_simd_helper() accepts only hex digits so we 
> > > >>>>>>>>> could
> > > >>>>>>>>> re-use it for UUID parsing. Let me check if the above idea of 
> > > >>>>>>>>> using
> > > >>>>>>>>> the return value works for us first.
> > > >>>>>>>>>
> > > >>>>>>>>
> > > >>>>>>>> That sounds reasonable.  My main concern was to keep the fast 
> > > >>>>>>>> path’s accepted
> > > >>>>>>>> input set identical to the scalar UUID parser.  Falling back 
> > > >>>>>>>> when the decoded
> > > >>>>>>>> length is not UUID_LEN, together with regression tests for 
> > > >>>>>>>> whitespace cases,
> > > >>>>>>>> should address that.
> > > >>>>>>>>
> > > >>>>>>>>>>
> > > >>>>>>>>>> Could you also add a few regression tests for invalid inputs 
> > > >>>>>>>>>> that contain
> > > >>>>>>>>>> whitespace inside otherwise fast-path-looking UUID strings?  
> > > >>>>>>>>>> For example:
> > > >>>>>>>>>>
> > > >>>>>>>>>> ---------------------------------------------------------------
> > > >>>>>>>>>>
> > > >>>>>>>>>> SELECT 'a0eebc99 9c0b4ef8bb6d6bb9bd380a11'::uuid;
> > > >>>>>>>>>> SELECT 'a0eebc999c0b4ef8bb6d6bb9bd380a1 '::uuid;
> > > >>>>>>>>>> SELECT '{a0eebc999c0b4ef8bb6d6bb9bd380a1 }'::uuid;
> > > >>>>>>>>>> SELECT 'a0eebc99-9c0b-4ef8-bb6d-6bb9bd380a1 '::uuid;
> > > >>>>>>>>>> ---------------------------------------------------------------
> > > >>>>>>>>>>
> > > >>>>>>>>>> These should continue to be rejected in the same way as the 
> > > >>>>>>>>>> scalar parser.
> > > >>>>>>>>>> Regards,
> > > >>>>>>>>>
> > > >>>>>>>>> Agreed.
> > > >>>>>>>>>
> > > >>>>>>>
> > > >>>>>>> I've attached the updated patch.
> > > >>>>>>>
> > > >>>>>>> Regards,
> > > >>>>>>>
> > > >>>>>>> --
> > > >>>>>>> Masahiko Sawada
> > > >>>>>>> Amazon Web Services: https://aws.amazon.com
> > > >>>>>>
> > > >>>>>> I noticed a few typos in the comments:
> > > >>>>>>
> > > >>>>>> src/backend/utils/adt/uuid.c
> > > >>>>>> line 56: “scalar implmentation” -> “scalar implementation”
> > > >>>>>> line 109: “swalled” -> “swallowed”
> > > >>>>>> line 110: “kepping” -> “keeping”
> > > >>>>>> line 118: “grammer” -> “grammar”
> > > >>>>>> line 119: “whitespaces” -> “whitespace”
> > > >>>>>>
> > > >>>>>> Could you fix them ?
> > > >>>>>
> > > >>>>> Oops, I fixed them and rechecked other places.
> > > >>>>>
> > > >>>>> I've attached the updated patch.
> > > >>>>>
> > > >>>>> Regards,
> > > >>>>>
> > > >>>>> --
> > > >>>>> Masahiko Sawada
> > > >>>>> Amazon Web Services: https://aws.amazon.com
> > > >>>>
> > > >>>> The code looks good to me now.  I only noticed one small typo in the
> > > >>>> commit trailer: Reviwed-by should be Reviewed-by.
> > > >>>>
> > > >>>> Otherwise, it looks good.  Thank you for fixing these issues.
> > > >>>>
> > > >>>
> > > >>> After spending more time on this patch, I find out two things:
> > > >>>
> > > >>> 1. USE_NO_SIMD doesn't work in uuid.c without including port/simd.h.
> > > >>> But including port/simd.h seems wrong as it doesn't use any SIMD
> > > >>> support functions.
> > > >>>
> > > >>> 2. hex_decode_safe() is faster than the current UUID parse
> > > >>> (isxdigit()+strtoul() approach) even without SIMD. I've created a
> > > >>> small benchmark test tool (attached as 0002 patch, not intended to be
> > > >>> pushed into the core), and measures UUID parsing performance of three
> > > >>> approaches: 'scalar' is the current string_to_uuid() that uses
> > > >>> isxdigit()+strtoul()), 'simd' uses hex_decode_safe() with SIMD, and
> > > >>> 'nosimd' uses hex_decode_safe() without SIMD, with different shapes of
> > > >>> UUIDs. Here are results:
> > > >>>
> > > >>> =# select path, shape, n_inputs, best_ms::numeric(10,3) from
> > > >>> uuid_parse_bench(100000, 5);
> > > >>> path  |      shape       | n_inputs | best_ms
> > > >>> --------+------------------+----------+---------
> > > >>> scalar | canonical        |   100000 |  22.661
> > > >>> simd   | canonical        |   100000 |   1.400
> > > >>> nosimd | canonical        |   100000 |   1.652
> > > >>> scalar | bare32           |   100000 |  15.932
> > > >>> simd   | bare32           |   100000 |   0.471
> > > >>> nosimd | bare32           |   100000 |   1.110
> > > >>> scalar | braced_canonical |   100000 |  17.330
> > > >>> simd   | braced_canonical |   100000 |   1.088
> > > >>> nosimd | braced_canonical |   100000 |   1.314
> > > >>> scalar | braced_bare32    |   100000 |  15.942
> > > >>> simd   | braced_bare32    |   100000 |   0.488
> > > >>> nosimd | braced_bare32    |   100000 |   1.141
> > > >>> scalar | dashed4          |   100000 |  16.185
> > > >>> simd   | dashed4          |   100000 |  16.493
> > > >>> nosimd | dashed4          |   100000 |  16.403
> > > >>> scalar | invalid_hex      |   100000 |   0.199
> > > >>> simd   | invalid_hex      |   100000 |   1.150
> > > >>> nosimd | invalid_hex      |   100000 |   0.385
> > > >>> (18 rows)
> > > >>>
> > > >>> Each of shape means:
> > > >>> - 'canonical': 8x-4x-4x-4x-12x, what uuid_out() emits
> > > >>> - 'bare32': 32 contiguous hex digits
> > > >>> - 'braced_canonical': {8x-4x-4x-4x-12x}
> > > >>> - 'braced_bare32': {32 hex digits}
> > > >>> - 'dashed4': dash after every group of 4
> > > >>> - 'invalid_hdx': canonical but with a invalid digit
> > > >>>
> > > >>> 'nosimd' is 10x~ faster than 'scalar' in most cases. All paths are
> > > >>> mostly the same in 'dashed4' and 'invalid_hex' cases because 'simd'
> > > >>> and 'nosimd' fall back to the 'scalar' case. According to these
> > > >>> results, my conclusion is that we can use hex_decode_safe() for
> > > >>> canonical forms and 32 contiguous hex forms anyway, and let
> > > >>> hex_decode_safe() choose whether to use SIMD. We would win in either
> > > >>> case. We still use the current scalar approach for uncommon UUID forms
> > > >>> and error reporting purposes.
> > > >>>
> > > >>> Regards,
> > > >>>
> > > >>> --
> > > >>> Masahiko Sawada
> > > >>> Amazon Web Services: https://aws.amazon.com
> > > >>> <v4-0001-Optimize-UUID-parse-using-SIMD.patch><v4-0002-uuid_parse_bench-module.patch>
> > > >>
> > > >> A few comments on v4.
> > > >>
> > > >> 1 - 0001
> > > >> ```
> > > >> +static void
> > > >> +string_to_uuid(const char *source, pg_uuid_t *uuid, Node *escontext)
> > > >> +{
> > > >> +       const char *body = source;
> > > >> +       size_t          len = strlen(source);
> > > >> ```
> > > >>
> > > >> I think it would be better to avoid strlen(). The old code processes 
> > > >> at most UUID_LEN (16) byte pairs, so it does not need to scan 
> > > >> arbitrarily far on malformed input. So, maybe we could use something 
> > > >> like strnlen(source, 39) instead.
> > > >
> > > > While strnlen(source, 39) works there, 39 is a magic number and it's
> > > > tied to the current format check logic. What is the benefit of using
> > > > strnlen(source, 39) instead? I'm not sure it warrants having the magic
> > > > number.
> > >
> > > It doesn't have to be exactly 39; 1024 (long enough) would also work, or 
> > > perhaps something based on UUID_LEN, such as UUID_LEN * 3. I think the 
> > > main point is to avoid unbounded scanning on malformed input.
> > >
> > > The old code did not have this issue because it only examined as much 
> > > input as needed based on UUID_LEN. The new fast path starts to use 
> > > strlen(), so this would be a new risk introduced by the optimization.
> >
> > I don't think the scan can be really unbounded. string_to_uuid()
> > receives a cstring, so by the time it is called the caller has already
> > walked or copied the whole string to produce it. So unless the
> > unbounded scan can be reached in some path I have overlooked, I'd
> > prefer to keep strlen() here. Happy to change it if you still think it
> > is worth it.
>
> After more thoughts, while I still don't think the scan can be
> unbounded, using strlen() would add an extra scan just to determine we
> use hex_decode_safe(). I'll change it to use strnlen() instead.

I've updated the patch accordingly. Please review it.

Regards,

-- 
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com
From 52e9030780fadcaf905062ceefecb4f87737ab9d Mon Sep 17 00:00:00 2001
From: Masahiko Sawada <[email protected]>
Date: Thu, 25 Jun 2026 10:03:44 -0700
Subject: [PATCH v5 1/2] Optimize UUID parse using SIMD.

Previously, string_to_uuid() decoded one byte at a time, calling
isxdigit() twice and strtoul() once for every pair of hexadecimal
digits.  That loop dominated the cost of uuid_in().

This commit adds a fast path for the two common shapes: a bare string
of 32 hexadecimal digits, and the canonical 8x-4x-4x-4x-12x form
(where "nx" means n hexadecimal digits), each optionally wrapped in
braces.  Both are compacted into 32 contiguous hexadecimal digits and
decoded with hex_decode_safe().  Any other shape, or any decoding
error, is handed off to the original scalar parser, now
string_to_uuid_scalar(), so the accepted grammar and the error
messages are unchanged.

hex_decode_safe() silently skips whitespace while the UUID grammar
does not, so a decode can succeed and still write fewer than UUID_LEN
bytes.  The fast path therefore treats a short result as a failure,
just like an error, and lets the scalar parser reject the input and
report the syntax error.

The fast path is deliberately not conditional on SIMD support.
hex_decode_safe() selects a vectorized or scalar implementation
itself, and even its scalar implementation is an order of magnitude
faster than decoding a byte at a time, so gating this on USE_NO_SIMD
would only penalize platforms that have neither SSE2 nor NEON.

Reviewed-by: Bharath Rupireddy <[email protected]>
Reviewed-by: Haibo Yan <[email protected]>
Discussion: https://postgr.es/m/cad21aocqer4uqu77q_yomnnzj7aveio5qzt+4hnzpm4wm-e...@mail.gmail.com
---
 src/backend/utils/adt/uuid.c       | 103 +++++++++++++++++++++++++++--
 src/test/regress/expected/uuid.out |  75 +++++++++++++++++++++
 src/test/regress/sql/uuid.sql      |  25 +++++++
 3 files changed, 198 insertions(+), 5 deletions(-)

diff --git a/src/backend/utils/adt/uuid.c b/src/backend/utils/adt/uuid.c
index 28e18940a9d..9a3b6715828 100644
--- a/src/backend/utils/adt/uuid.c
+++ b/src/backend/utils/adt/uuid.c
@@ -19,7 +19,9 @@
 #include "common/hashfn.h"
 #include "lib/hyperloglog.h"
 #include "libpq/pqformat.h"
+#include "nodes/miscnodes.h"
 #include "port/pg_bswap.h"
+#include "utils/builtins.h"
 #include "utils/fmgrprotos.h"
 #include "utils/guc.h"
 #include "utils/skipsupport.h"
@@ -139,13 +141,13 @@ uuid_out(PG_FUNCTION_ARGS)
 }
 
 /*
- * We allow UUIDs as a series of 32 hexadecimal digits with an optional dash
- * after each group of 4 hexadecimal digits, and optionally surrounded by {}.
- * (The canonical format 8x-4x-4x-4x-12x, where "nx" means n hexadecimal
- * digits, is the only one used for output.)
+ * Reference implementation of the UUID grammar, parsing one character at a
+ * time.  string_to_uuid() recognizes the common shapes more cheaply and
+ * defers to this function for everything else, so this is also the only
+ * place that reports a syntax error.
  */
 static void
-string_to_uuid(const char *source, pg_uuid_t *uuid, Node *escontext)
+string_to_uuid_scalar(const char *source, pg_uuid_t *uuid, Node *escontext)
 {
 	const char *src = source;
 	bool		braces = false;
@@ -194,6 +196,97 @@ syntax_error:
 					"uuid", source)));
 }
 
+/*
+ * We allow UUIDs as a series of 32 hexadecimal digits with an optional dash
+ * after each group of 4 hexadecimal digits, and optionally surrounded by {}.
+ * (The canonical format 8x-4x-4x-4x-12x, where "nx" means n hexadecimal
+ * digits, is the only one used for output.)
+ *
+ * The two common shapes -- a bare string of 32 hexadecimal digits and the
+ * canonical form, each optionally wrapped in braces -- are compacted into 32
+ * contiguous hex digits and decoded with hex_decode_safe(), which is much
+ * faster than the character-at-a-time loop.  Any other shape, or any decoding
+ * error, is handed off to string_to_uuid_scalar() so that the accepted
+ * grammar and the error messages are unchanged.
+ *
+ * Note that this fast path is not conditional on SIMD support:
+ * hex_decode_safe() picks a vectorized or scalar implementation itself, and
+ * even its scalar implementation is far faster than string_to_uuid_scalar().
+ */
+static void
+string_to_uuid(const char *source, pg_uuid_t *uuid, Node *escontext)
+{
+	const char *body = source;
+	const char *hexsrc = NULL;
+	char		hexbuf[32];
+	uint64		written;
+	size_t		len;
+	ErrorSaveContext esctx = {T_ErrorSaveContext};
+
+	/*
+	 * Measure the input only far enough to classify its shape.  The bound
+	 * must exceed the longest shape handled here, the braced canonical form
+	 * at 38 characters: strnlen() returns the bound for anything at least
+	 * that long, so stopping at an accepted length would accept a longer
+	 * string that merely starts with a valid UUID.
+	 */
+	len = strnlen(source, 64);
+
+	/* Strip one optional surrounding brace pair */
+	if (len >= 2 && source[0] == '{' && source[len - 1] == '}')
+	{
+		body = source + 1;
+		len -= 2;
+	}
+
+	if (len == 32)
+	{
+		/*
+		 * Body is already 32 contiguous hex digits -- decode straight from
+		 * the input. hex_decode_safe() reads exactly body[0..31], so it never
+		 * touches the trailing NULL or '}'.
+		 */
+		hexsrc = body;
+	}
+	else if (len == 36 && body[8] == '-' && body[13] == '-' &&
+			 body[18] == '-' && body[23] == '-')
+	{
+		/*
+		 * Canonical 8x-4x-4x-4x-12x form; compact them into hexbuf with
+		 * fixed-offset copies, dropping the dashes.
+		 */
+		memcpy(&hexbuf[0], &body[0], 8);
+		memcpy(&hexbuf[8], &body[9], 4);
+		memcpy(&hexbuf[12], &body[14], 4);
+		memcpy(&hexbuf[16], &body[19], 4);
+		memcpy(&hexbuf[20], &body[24], 12);
+		hexsrc = hexbuf;
+	}
+
+	if (hexsrc == NULL)
+	{
+		/* Uncommon shape; let the general parse handle it */
+		string_to_uuid_scalar(source, uuid, escontext);
+		return;
+	}
+
+	/*
+	 * Decode the UUID hex data using our hex decoder that is SIMD-aware. We
+	 * give it a private error context so that a decode failure is swallowed
+	 * here and reported by the scalar path instead, keeping the error message
+	 * identical.
+	 */
+	written = hex_decode_safe(hexsrc, 32, (char *) uuid->data, (Node *) &esctx);
+
+	/*
+	 * Fall back to the scalar path on any error. We must also reject a short
+	 * result: hex_decode_safe() skips whitespace, so it can succeed yet write
+	 * fewer than UUID_LEN bytes, whereas the UUID grammar forbids whitespace.
+	 */
+	if (esctx.error_occurred || written != UUID_LEN)
+		string_to_uuid_scalar(source, uuid, escontext);
+}
+
 Datum
 uuid_recv(PG_FUNCTION_ARGS)
 {
diff --git a/src/test/regress/expected/uuid.out b/src/test/regress/expected/uuid.out
index d542eb14b26..6247bae8574 100644
--- a/src/test/regress/expected/uuid.out
+++ b/src/test/regress/expected/uuid.out
@@ -375,5 +375,80 @@ SELECT v = v::bytea::uuid as matched FROM gen_random_uuid() v;
  t
 (1 row)
 
+-- Test UUID shapes that the parser uses the SIMD path.
+SELECT '5b35380a-7143-4912-9b55-f322699c6770'::uuid;
+                 uuid                 
+--------------------------------------
+ 5b35380a-7143-4912-9b55-f322699c6770
+(1 row)
+
+SELECT '{5b35380a-7143-4912-9b55-f322699c6770}'::uuid;
+                 uuid                 
+--------------------------------------
+ 5b35380a-7143-4912-9b55-f322699c6770
+(1 row)
+
+SELECT '5b35380a714349129b55f322699c6770'::uuid;
+                 uuid                 
+--------------------------------------
+ 5b35380a-7143-4912-9b55-f322699c6770
+(1 row)
+
+SELECT '{5b35380a714349129b55f322699c6770}'::uuid;
+                 uuid                 
+--------------------------------------
+ 5b35380a-7143-4912-9b55-f322699c6770
+(1 row)
+
+-- Test if the UUID parser using SIMD optimization correctly rejects invalid UUID
+-- string format.
+SELECT '5b35380a714349129b55f32  99c6770'::uuid;
+ERROR:  invalid input syntax for type uuid: "5b35380a714349129b55f32  99c6770"
+LINE 1: SELECT '5b35380a714349129b55f32  99c6770'::uuid;
+               ^
+SELECT '5b35380a-7143-4912-9b55-f322699c67  '::uuid;
+ERROR:  invalid input syntax for type uuid: "5b35380a-7143-4912-9b55-f322699c67  "
+LINE 1: SELECT '5b35380a-7143-4912-9b55-f322699c67  '::uuid;
+               ^
+SELECT '  35380a-7143-4912-9b55-f322699c6770'::uuid;
+ERROR:  invalid input syntax for type uuid: "  35380a-7143-4912-9b55-f322699c6770"
+LINE 1: SELECT '  35380a-7143-4912-9b55-f322699c6770'::uuid;
+               ^
+SELECT 'AZ35380a-7143-4912-9b55-f322699c6770'::uuid;
+ERROR:  invalid input syntax for type uuid: "AZ35380a-7143-4912-9b55-f322699c6770"
+LINE 1: SELECT 'AZ35380a-7143-4912-9b55-f322699c6770'::uuid;
+               ^
+SELECT '{AZ35380a-7143-4912-9b55-f322699c6770}'::uuid;
+ERROR:  invalid input syntax for type uuid: "{AZ35380a-7143-4912-9b55-f322699c6770}"
+LINE 1: SELECT '{AZ35380a-7143-4912-9b55-f322699c6770}'::uuid;
+               ^
+SELECT '{AZ35380a714349129b55f322699c6770}'::uuid;
+ERROR:  invalid input syntax for type uuid: "{AZ35380a714349129b55f322699c6770}"
+LINE 1: SELECT '{AZ35380a714349129b55f322699c6770}'::uuid;
+               ^
+SELECT '{AZ35380a714349129b55f322699c67  }'::uuid;
+ERROR:  invalid input syntax for type uuid: "{AZ35380a714349129b55f322699c67  }"
+LINE 1: SELECT '{AZ35380a714349129b55f322699c67  }'::uuid;
+               ^
+-- The parser only measures the input far enough to classify its shape.  If it
+-- stopped measuring at one of the accepted lengths, a longer string that
+-- merely starts with a valid UUID would look like that UUID and be accepted
+-- with the rest silently ignored, so check that trailing data is rejected.
+SELECT '5b35380a714349129b55f322699c6770TRAILING'::uuid;
+ERROR:  invalid input syntax for type uuid: "5b35380a714349129b55f322699c6770TRAILING"
+LINE 1: SELECT '5b35380a714349129b55f322699c6770TRAILING'::uuid;
+               ^
+SELECT '{5b35380a714349129b55f322699c6770}TRAILING'::uuid;
+ERROR:  invalid input syntax for type uuid: "{5b35380a714349129b55f322699c6770}TRAILING"
+LINE 1: SELECT '{5b35380a714349129b55f322699c6770}TRAILING'::uuid;
+               ^
+SELECT '5b35380a-7143-4912-9b55-f322699c6770TRAILING'::uuid;
+ERROR:  invalid input syntax for type uuid: "5b35380a-7143-4912-9b55-f322699c6770TRAILING"
+LINE 1: SELECT '5b35380a-7143-4912-9b55-f322699c6770TRAILING'::uuid;
+               ^
+SELECT '{5b35380a-7143-4912-9b55-f322699c6770}TRAILING'::uuid;
+ERROR:  invalid input syntax for type uuid: "{5b35380a-7143-4912-9b55-f322699c6770}TRAILING"
+LINE 1: SELECT '{5b35380a-7143-4912-9b55-f322699c6770}TRAILING'::uui...
+               ^
 -- clean up
 DROP TABLE guid1, guid2, guid3 CASCADE;
diff --git a/src/test/regress/sql/uuid.sql b/src/test/regress/sql/uuid.sql
index 54f0d8f8255..d13e8c21921 100644
--- a/src/test/regress/sql/uuid.sql
+++ b/src/test/regress/sql/uuid.sql
@@ -178,5 +178,30 @@ SELECT '\x019a2f859ced7225b99d9c55044a2563'::bytea::uuid;
 SELECT '\x1234567890abcdef'::bytea::uuid; -- error
 SELECT v = v::bytea::uuid as matched FROM gen_random_uuid() v;
 
+-- Test UUID shapes that the parser uses the SIMD path.
+SELECT '5b35380a-7143-4912-9b55-f322699c6770'::uuid;
+SELECT '{5b35380a-7143-4912-9b55-f322699c6770}'::uuid;
+SELECT '5b35380a714349129b55f322699c6770'::uuid;
+SELECT '{5b35380a714349129b55f322699c6770}'::uuid;
+
+-- Test if the UUID parser using SIMD optimization correctly rejects invalid UUID
+-- string format.
+SELECT '5b35380a714349129b55f32  99c6770'::uuid;
+SELECT '5b35380a-7143-4912-9b55-f322699c67  '::uuid;
+SELECT '  35380a-7143-4912-9b55-f322699c6770'::uuid;
+SELECT 'AZ35380a-7143-4912-9b55-f322699c6770'::uuid;
+SELECT '{AZ35380a-7143-4912-9b55-f322699c6770}'::uuid;
+SELECT '{AZ35380a714349129b55f322699c6770}'::uuid;
+SELECT '{AZ35380a714349129b55f322699c67  }'::uuid;
+
+-- The parser only measures the input far enough to classify its shape.  If it
+-- stopped measuring at one of the accepted lengths, a longer string that
+-- merely starts with a valid UUID would look like that UUID and be accepted
+-- with the rest silently ignored, so check that trailing data is rejected.
+SELECT '5b35380a714349129b55f322699c6770TRAILING'::uuid;
+SELECT '{5b35380a714349129b55f322699c6770}TRAILING'::uuid;
+SELECT '5b35380a-7143-4912-9b55-f322699c6770TRAILING'::uuid;
+SELECT '{5b35380a-7143-4912-9b55-f322699c6770}TRAILING'::uuid;
+
 -- clean up
 DROP TABLE guid1, guid2, guid3 CASCADE;
-- 
2.55.0

Reply via email to