Content by mariko wakabayashi (1)

How to evaluate LLMs before production

Mariko Wakabayashi shares practical lessons from evaluating an LLM-based system to reduce false positives in GitHub secret scanning, focusing on how to define success metrics, keep offline evaluation close to production, and use error analysis and human review to manage security risk.
News

End of content

Rejoining the server...

Rejoin failed... trying again in seconds.

Failed to rejoin.
Please retry or reload the page.

The session has been paused by the server.

Failed to resume the session.
Please reload the page.