ROC-AUC at 0.97, Precision at 9%: The Base-Rate Math
Work through a 1%-positive detector, turn thresholds into ROC and precision-recall points, and choose an operating point for a 30-alert review budget.
Tag
9articleswith this tag.
Work through a 1%-positive detector, turn thresholds into ROC and precision-recall points, and choose an operating point for a 30-alert review budget.
OpenAI and AWS report that GPT-5.6 Terra completed successful Terminal-Bench tasks in Kiro at roughly 82% lower cost, but repository work adds review and repair.
Read Anthropic's August 2026 Risk Report by separating its coverage date, company ratings, disclosed failures, redactions, and external review.
Visual Studio 18.10 Insiders can connect local and hosted models to Agent mode, but one hands-on test shows that endpoint availability does not imply reliable coding work.
Thomson Reuters released a proprietary model for CoCounsel and open weights for a smaller version. Here is what buyers can verify—and what remains a vendor claim.
Hugging Face has passed three million models, but likes and downloads reveal attention rather than license fit, provenance, hardware cost, or task quality.
Gemini 3.7 Flash prices rise in January 2027, while the cheapest thinking level still depends on acceptance, retries, latency, and fallback cost.
DeepSeek has added image input to its fast API line, but the preview does not yet establish reliable performance across screenshots, charts, and documents.
Grok 4.6 combines low headline pricing with strong reported benchmarks, but retries, failed tasks, latency, and tool charges can reverse the apparent advantage.