this post was submitted on 28 Sep 2026
77 points (95.3% liked)

Technology

88365 readers
4581 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] RumRunningDevil@lemmy.zip 5 points 1 day ago (2 children)

To quote the research paper conclusion...

Conclusion We evaluate the impact of context files on coding agent performance for four common coding agents on SWE-BENCH and the novel CTXBENCH, built from recent GitHub issues and less popular repositories containing developer-written context files. We find that all context files consistently increase the cost and number of steps required to complete tasks. LLM-generated context files have a marginal negative effect on task success rates, while developer-written ones provide a marginal performance gain, neither statistically significant. Our trace analyses show that instructions in context files are generally followed and lead to more test- ing and broader exploration; however, they do not function as effective repository overviews. Over- all, our results suggest that context files don’t improve coding agent performance, and should only contain specific additional instructions beyond what is already available in the codebase. This high- lights a concrete gap between current agent-developer recommendations and observed outcomes, and motivates future work on principled ways to automatically generate concise, task-relevant guid- ance for coding agents.

That sounds like "Agents.MD doesn't work" to me.

Am I missing something?

[–] brian@programming.dev 1 points 23 hours ago

your research paper:

llm generated agent.md files are harmful. handwritten ones on average don't have an effect on ability to complete a task

thread article:

agents keep signing off on commits wrong. we added it to the agents.md and now they don't

benchmarks notoriously don't measure the likelihood that the results would be merged by maintainers, just that they passed a test

the old style "repo map" agents.md files that every tool used to create is no longer useful, but of course things that would otherwise be repeated prompts are still good. I cut the ones at work down from hundreds of lines each to tens. without those last few lines the amount of iteration I do - both with agent and when reviewing PRs - goes up significantly

[–] Feathercrown@lemmy.world 1 points 1 day ago (1 children)

Agents.md doesn't exist to make sure generated code functions, it's there to make sure it's written well.

[–] RumRunningDevil@lemmy.zip 2 points 1 day ago (2 children)

What is the difference between the models writing code "well" and their performance in this context? Are we referring to readability?

Genuine question. If we use agents to read, edit, and review code, why do we care about readability? That's a human constraint. Unless attempting to do those three is not effective and thus requires human attention to correct issues which would justify readable code. If that's the case; why use the agent to edit the code in the first place?

[–] Feathercrown@lemmy.world 1 points 23 hours ago* (last edited 23 hours ago)

"Writing code well" includes several relavant things:

  • Security implications
  • Stability and edge case handling
  • Performance
  • Quality of output (ie. for UI/UX)
  • Readability (Agents need to read to make changes too! Violating DRY/SOLID/etc. could still cause issues!)

As far as I'm aware you do currently still need a human to ensure stuff like this is followed. People use agents because they don't care about the above, or because they can get close enough and intervene to fix any issues that appear.

[–] hirihit640@sh.itjust.works 1 points 1 day ago* (last edited 1 day ago) (1 children)

If that's the case; why use the agent to edit the code in the first place?

Agents for generation, humans for review

[–] RumRunningDevil@lemmy.zip 2 points 1 day ago* (last edited 1 day ago) (1 children)

One, that sounds truly miserable.

Two, there is a decent body of evidence to suggest that this method does not actually speed up development unless you go full lights out software factory which... Why would you want to do that?

https://ide.mit.edu/insights/ai-productivity-and-roi/

EDIT: As an aside this actually reminds me of the Xerox park study into efficiency gains for keyboard heavy workflows vice mouse heavy workflows. Keyboards were perceived as faster by subjects but when actually measured the mouse was faster.

[–] Feathercrown@lemmy.world 2 points 23 hours ago (1 children)

That's just because I need to stop and appreciate my efficiency for 10 seconds every time it saves me 10 seconds :-)

[–] RumRunningDevil@lemmy.zip 2 points 21 hours ago* (last edited 21 hours ago) (1 children)

You joke but that's basically the conclusion of both the newer agent studies and the older Xerox park study on peripheral use. Language and logic are stored in an easier to access location in the brain than the kinds of skills used in architecture and planning.

To be clear, I'm no rube. I recognize that these are powerful tools. But I suppose my take here is that if an LLM is necessary to get rid of a lot of the boilerplate and setup for a task that indicates we should explore how we're doing the task.

What I want to see is using these tools not to delegate our problem solving but finding ways to enhance and accelerate it and I don't think writing spec sheets is the answer.

[–] Feathercrown@lemmy.world 2 points 20 hours ago* (last edited 20 hours ago)

I think you're correct, the amount of rigor in your prompts seems orthogonal from whether you're using them for architecture or implementation.