While this may be true that MOST projects don't need the complexity that the internet proselytizes, if you want a job in tech, you'll have to learn the teachings of the internet and what the "industry" agrees are "best practices".
This feels so dishonest. If the vulnerabilities are a needle in the haystack. Mythos was just given the haystack and told to find the needle while the authors pointed to a spot in the haystack and told their LLM to try looking around there. That's not even close to being the same.
so how would you eval your own claude.md? Each context is unique to the project, team, and personal root claude.md. Do you just take given task and ask it to redo the same one over and over again against a known solution? Do you just keep using it and "feel" whether or not it's working? How is that different from what everyone is already doing?
The review eval tests language, activation etc of skills. I guess you could move it all to a skill quick and then run an eval on that if using Tessl. This checks if the way you write the instructions etc are being well understood by the agent