I was deep in thought about how much of a model hallucination is RLHF-based assumption in the absence of training data or context. Pretraining gives the model the ability to confabulate. Post-training often influences whether it chooses to confabulate rather than say "I don't know."
A base language model is trained to predict plausible continuations. If the evidence needed to answer is ...
Software has borrowed a surprising amount from manufacturing. Not from manufacturing in the sense of “let’s put programmers on an assembly line.” Thankfully, we’ve tried enough variations of that idea already.
I mean something more interesting.
Some of the most influential ideas in modern software development came from looking at how Toyota transformed manufacturing. Lean Softwa...
Internet-wide scanners are a double-edged tool. Attackers use them to find exposed systems; defenders can use the same capability to discover what they have accidentally exposed. In 2026, government advisories have repeatedly pointed to scanning services as part of the attacker's workflow, w...
Originally published at erikhill.dev. The numbers below are checked against the repository they come from.
This is a finding about a measurement rule, not about a model....
Originally published at erikhill.dev. The numbers below are checked against the repository they come from.
This is a finding about my own harness. The suspect is th...
$87,000 vs roughly $1,000. Those were our first estimates for processing a corpus of 6.5 million semantic atoms with Sonnet or training and running a smaller model ourselves. I was ready to build the smaller model.
Our job was narrow. We needed to read financial text and record who said what, what they claimed, and how certain the source was. A wrong speaker or an invented certainty wou...