t18: Feature with changes -- GPT-6.1 Sol vs Claude Opus 5.5


Add tags to a Python todo app, then handle three change requests that arrived later.
Opus: 74.3 s. Sol: 194.8 s. Sol took 2.6 times as long.
Opus (s)Sol (s)
Hidden tests: Opus 8 of 8, Sol 8 of 8.
Blind judge pick: Opus (small edge).
"Both solutions perfectly implement the requirements and edge cases. Solution A [Opus] is slightly better structured, utilizing a helper function for argparse to avoid duplication, organizing tests into logical classes, and using a full snapshot of the legacy code for robust compatibility testing."
"An outstanding solution with clean implementation, robust validation, and excellent compatibility testing using a legacy snapshot."
"An excellent, complete, and well-tested solution that perfectly meets all requirements and edge cases."