All Projects
Each project ships with real results — actual numbers, not just screenshots. Built progressively as the roadmap unfolds.
AI Safety · LLM Evaluation · Moderation
phi3:mini — the smallest model — outperformed both larger models on safety.
Prompt Engineering · Robustness · Security
CoT pushed Q&A accuracy to 100% — but scored the lowest robustness of all strategies.
Prompt Optimization · LLM-as-judge · Evolutionary Search
89% token reduction with a higher score. The optimizer found it automatically.
Retrieval · Embeddings · Evaluation
Benchmarking chunk size & top-k retrieval across 4 RAGAS metrics.