reef-eval added to PyPI
Evaluation infrastructure for self-evolving agents on the Harbor task standard: autoresearch and task streams, tamper-proof scoring, budgets on four axes, and resumable sweeps.
Evaluation infrastructure for self-evolving agents, on the Harbor task standard. An agent self-evolves when something it learned during a run persists past it: memory, a skill library, an evolved h… [+10173 chars]