All Projects

Work

Each project ships with real results — actual numbers, not just screenshots. Built progressively as the roadmap unfolds.

01

Toxicity Stress-Test & Moderation Pipeline

AI Safety · LLM Evaluation · Moderation

phi3:mini — the smallest model — outperformed both larger models on safety.

Live
PythonDetoxifyOllamaStreamlit
02

Prompt Robustness Evaluation Harness

Prompt Engineering · Robustness · Security

CoT pushed Q&A accuracy to 100% — but scored the lowest robustness of all strategies.

Live
PythonOllamaPlotly DashROUGE
03

Adaptive Prompt Optimization Pipeline

Prompt Optimization · LLM-as-judge · Evolutionary Search

89% token reduction with a higher score. The optimizer found it automatically.

Live
PythonOllamaPlotly DashDSPy-inspired
04

RAG Evaluation Pipeline with RAGAS

Retrieval · Embeddings · Evaluation

Benchmarking chunk size & top-k retrieval across 4 RAGAS metrics.

Planned
PythonFAISSsentence-transformersRAGAS