Content by Julia Kasper (2)

What 50,000 Runs of a 5-Line Eval Taught Us

Julia Kasper (VS Code Eval Team) breaks down what an ultra-simple “write HELLO.txt” agent eval revealed after 50,000+ runs: models vary widely in tool-call discipline, planning overhead, and output-token cost, and those differences matter for latency, billing, and automatic model selection in VS Code.
News
Julia Kasper explains the “coding harness” that powers GitHub Copilot’s agent experience in VS Code: how prompts and workspace context are assembled, how tools are exposed and executed, and how the agent loop is controlled and evaluated as models and providers change.
News

End of content

Rejoining the server...

Rejoin failed... trying again in seconds.

Failed to rejoin.
Please retry or reload the page.

The session has been paused by the server.

Failed to resume the session.
Please reload the page.