The impacts of coding agents on design, the future of software, dynamic media, interface invention, … apps-and-programming-two-accidental-tyrannies

RL prefers strategies with high expected reward and no safety margin (Libratus overbets, Diplodocus abandoning home centers); human experts are worst-case-weighted. Expert players said the bot's strategy would improve their average but they "hate" the variance.

Link to original

https://danluu.com/agentic-testing/
tldr:

  • agents are really bad at testing
  • not having any skills or frameworks for testing performs generally better than using popular skills or frameworks (not telling the agent to do useless work performs better and is cheaper than telling it to do useless work)
    • naming a technique doesn’t change agent behaviour; setting up a concrete test structure and nudging after seeing output does. Skills should modify the model’s default behaviour, not teach it from scratch.