I have a feeling that their test case is also a bit flawed. Trying to get index_value instead of index value is something I can imagine happening, and asking an LLM to ‘fix this but give no explanation’ is asking for a bad solution.
I think they are still correct in the assumption that output becomes worse, though

Yeah, if only QA vere not the first ‘replaced’ by AI 😠