Tag: Tool Calling
-

Thinking vs. No-Thinking in Local LLM Agents: Five Models, Two Trials, and No Universal Winner
Five local model artifacts tested twice with and without thinking, followed by GSM8K, MMLU, and IFEval: quality, safety, latency, and deployment trade-offs.
-

Six Local Models on Tool-Eval-Bench: Tool Calling, Safety, and State
Six quantized local models across 69 deterministic tool-calling scenarios, with observed quality, safety, state handling, latency, token use, and explicit limitations.
Read more about Six Local Models on Tool-Eval-Bench: Tool Calling, Safety, and State