
Text Mining Specialist
UW Libraries & eScience Institute
Mar 2026 → PresentCurrent
- Ingested 50,000+ of a 400,000-article target into Open IRE, an automated pipeline I built to collect, rights-classify, and preserve UW-authored articles, owning the database schema and system architecture.
- Recovered 6,800+ missing documents by building a custom exponential-backoff retry middleware and fixing a self-duplicate dedup bug behind a scraping pipeline failure, validated with a 295-test end-to-end suite.
- Improved recall@5 by 14pp on a 500-query eval set by redesigning RAG chunking from fixed 512-token windows to a sentence-anchored sliding window at 20% overlap, then shipped it as a Haystack service that cut manual research query time roughly 30%.
- Instructed at the AI in Practice Summer Institute, teaching applied AI workflows to researchers and library staff.
- Python
- RAG
- Haystack
- NLP
- PostgreSQL
- Pytest

