Because I almost never go full harness, I forget that harnesses are a fine way to produce code. Often, I sit much closer to the zero-shot side of the spectrum, with a few iterations and myself as the validator.
This got me thinking about the degree to which frontier model developers privilege our harness-produced code when generating their synthetic data. To the degree that I am willing to pay for tokens, I am, in some sense, vouching for the potential training value of the generated code. Even if it’s wrong, it may be better than random, in that it is text someone wanted. Path dependence and all of that. Producing idiosyncratic generators seems a lot more valuable than slurping up and training on pure state.
.- .-.. .-.. / ..-. --- --- - . .-. ... / .- .-. . / .-- .-. --- -. --. / ... --- -- . / .- .-. . / ..- ... . ..-. ..- .-.. FRIAM Applied Complexity Group listserv Fridays 9a-12p Friday St. Johns Cafe / Thursdays 9a-12p Zoom https://bit.ly/virtualfriam to (un)subscribe http://redfish.com/mailman/listinfo/friam_redfish.com FRIAM-COMIC http://friam-comic.blogspot.com/ archives: 5/2017 thru present https://redfish.com/pipermail/friam_redfish.com/ 1/2003 thru 6/2021 http://friam.383.s1.nabble.com/
