- Designed evaluation loops for retrieval quality, edge-case handling, and safe-response behavior.
- Turned observed assistant failures into retrieval tuning and knowledge-base updates.
- Defined rollout-readiness checks for everyday IT and HR support scenarios.
Case study · AI support and evaluation
AskScout: evaluating an internal IT and HR support assistant
Quality and evaluation work for AskScout, the internal IT and HR AI support assistant, so answers stayed accurate and trustworthy before rollout.
Repeatable evaluation so an internal support assistant could be trusted
Project lead for evaluation and QA
Evaluation and QA workstream for AskScout, the internal IT and HR support assistant, coordinated with IT leadership and knowledge owners.
- IT support teams fielding everyday requests
- HR support teams answering routine employee questions
- Internal users who needed accurate, reliable answers
- Test retrieval and safe-response behavior before rollout rather than after trust was lost.
- Tie failed responses back to knowledge-base updates instead of ad hoc prompt tweaks.
- Build a repeatable evaluation posture the assistant could keep inheriting as it grew.
Make sure an internal AI assistant for employee IT and HR requests gave accurate, trustworthy answers before it went wide.
- Evaluation and test loops focused on retrieval quality, edge-case behavior, and safe responses.
- An iteration process connecting observed failures to retrieval tuning and knowledge-base updates.
- Rollout-readiness checks for everyday IT and HR support scenarios.
- Built evaluation loops covering retrieval quality, edge cases, and safe-response behavior for support workflows.
- Tied failed answers back to retrieval tuning and knowledge-base updates instead of ad hoc prompt tweaks.
- Left a repeatable evaluation posture for future improvements to the assistant.
Introduce a formal scorecard to track retrieval accuracy and safe behavior across releases.