Open-weight LLMs caught up to GPT-5 and Gemini 3 Pro in 2026. We ranked 6 worth running in production — by reasoning, speed, multimodal, and cost-per-token — with the tradeoffs to know before you deploy.
Hendrix Liu · July 29, 2026

A comprehensive comparison of Claude 3.5 Haiku and Claude 3.5 Sonnet covering benchmarks, speed, cost, and when to choose each model for production workloads.
Hendrix Liu · July 29, 2026Helicone alternatives need to replace both the observability layer and the gateway. Compare 10 platforms on what each covers and what it costs.
Hendrix Liu · July 27, 2026
I built a RAG tutor that answers only from my course PDFs. Here are the per-step evals that made it trustworthy, and what they caught.
Amy Kodama · July 23, 2026