The impacts of coding agents on design, the future of software, dynamic media, interface invention, … apps-and-programming-two-accidental-tyrannies
Link to originalRL prefers strategies with high expected reward and no safety margin (Libratus overbets, Diplodocus abandoning home centers); human experts are worst-case-weighted. Expert players said the bot's strategy would improve their average but they "hate" the variance.
https://danluu.com/agentic-testing/
tldr:
- agents are really bad at testing
- not having any skills or frameworks for testing performs generally better than using popular skills or frameworks (not telling the agent to do useless work performs better and is cheaper than telling it to do useless work)
-
- naming a technique doesn’t change agent behaviour; setting up a concrete test structure and nudging after seeing output does. Skills should modify the model’s default behaviour, not teach it from scratch.