Tag: AI Benchmarking
-

Thinking vs. No-Thinking in Local LLM Agents: Five Models, Two Trials, and No Universal Winner
Five local model artifacts tested twice with and without thinking, followed by GSM8K, MMLU, and IFEval: quality, safety, latency, and deployment trade-offs.
-

Six Local Models on Tool-Eval-Bench: Tool Calling, Safety, and State
Six quantized local models across 69 deterministic tool-calling scenarios, with observed quality, safety, state handling, latency, token use, and explicit limitations.
Read more about Six Local Models on Tool-Eval-Bench: Tool Calling, Safety, and State
-

Eight Quantized Local Models on Aider Polyglot: An Operational Snapshot
Eight quantized local models across 225 Aider Polyglot tasks, with observed Pass@2, protocol failures, wall time, token use, and explicit serving confounders.
Read more about Eight Quantized Local Models on Aider Polyglot: An Operational Snapshot
-

Debugging Local LLM Benchmarks: Aider, llama.cpp, and Runaway Output
A practical account of building a fair local-agent benchmark, debugging reasoning-model serving, fixing greedy-decoding loops, and adding an exact-repetition circuit breaker to llama.cpp.
Read more about Debugging Local LLM Benchmarks: Aider, llama.cpp, and Runaway Output