Does AGENTS.md Actually Work? The First Empirical Study
The first empirical study evaluating AGENTS.md effectiveness has been published. We analyze its impact on coding agent success rates and inference costs.
Tags
4 posts
The first empirical study evaluating AGENTS.md effectiveness has been published. We analyze its impact on coding agent success rates and inference costs.
Verbalized Sampling tackles mode collapse after alignment by prompting models to verbalize probability distributions, achieving 1.6-2.1x diversity gains without retraining
Experimental results and statistical analysis of 225 evaluations using LLM-based Semantic Similarity Rating. Validated high reliability with ICC 0.83 and visualizations.
Does giving an AI agent a gender or persona change performance? Drawing on 120+ psychology and NLP studies, we unpack expert personas, emotion, and role assignment — with design strategies by task type.