Agents and coding models
Compare model behavior, agent reliability, tool use and the systems that evaluate them.
What we currently think
Model rankings depend on the work, constraints and evaluator. A useful comparison preserves failures and tests transfer beyond one prompt set.
Start with your question
- Which model fits this kind of work?
- Does the evaluation measure the property we care about?
- What breaks when the agent operates a real system?
Start here
Kimi K3 vs GLM-5.2 vs GPT-5.6 Sol: Visual Benchmark
Eight frontier-model routes explain one product, revise their work, and turn evidence into decisions. Compare the finished artifacts, costs, failures, and limits.
GPT-5.6 Luna vs Terra vs Sol: 294 Codex Runs Explained
We compared GPT-5.6 Luna, Terra, and Sol across Codex reasoning levels using 294 ChatGPT-subscription runs. See why medium is the practical starting point, why no model won, and when higher reasoning may be worth the cost.
How We Used SkillOpt to Stress-Test Codex Skills
We adapted the SkillOpt idea into a practical test loop for Codex skills: synthetic edge cases, narrow edits, and fresh unseen score streaks before we trusted an improvement.
Continue exploring
LFM2.5-2.6B Training Explained - SFT, Specialist Teachers, MOPD and Agentic RL
How Liquid AI trained LFM2.5-2.6B for local agents using two-stage SFT, domain-specialist teachers, multi-domain on-policy distillation and reinforcement learning inside real agent harnesses.
Token Saver with Codex: A Smaller, Safer Way to Read Long PDFs
We tested whether Codex could answer focused questions from long PDFs without receiving every page. See what worked, what did not, and where the approach fits in an AI video workflow.
Nanbeige4.2-3B vs Qwen3-VL-4B on an RTX 3090 Ti
We ran 25 fixed arithmetic, planning, tool-call, and transformation tasks on Nanbeige4.2-3B and Qwen3-VL-4B. Here are the scores, latency, VRAM use, compatibility failure, and limits of this local agent benchmark.
Function Calling and MCP First Principles
A plain-language walkthrough of how function calling and MCP fit together in a production agent loop.
Mining Claude Code and Codex Logs Into a Knowledge Base
A practical guide to turning agent transcripts into reusable engineering memory without dumping every JSONL file into a vector database.
OpenAI Codex CLI /goal Command: What It Does and How to Enable It
A source-inspected guide to the experimental OpenAI Codex CLI /goal command: what it stores, how it lets Codex continue work, how to enable the goals feature today, and why Codex Desktop users should treat it as CLI-only for now.
Running OpenAI Privacy Filter on an M2 MacBook Pro - 52-Case Benchmark
Can OpenAI's new privacy-filter token classifier redact secrets and PII on a 16 GB M2 MacBook before they reach a cloud LLM? We ran 52 test cases covering API keys, PEM blocks, JWTs, names across six cultures, addresses, and decoy prose. Load time, latency, hit rates, and the AWS-key miss that matters.
Hardening Agents in Production - Locking Down the Attack Surface
A live conference-note draft on why agents are already in production, why autonomous workflows create a larger attack surface than most teams expect, and what a real security control plane must do.
Systemic Decision Rot - Why AI Governance Must Move Beyond Model Safety
A conference-note essay on why AI risk is not just about hallucinating models, but about how organizations transform AI outputs into summaries, recommendations, and institutional decisions without enough verification.
The Agent-to-Agent Internet - Evaluation Arenas, Algorithmic Governance, and the Dark Web of AI
A strategy memo on what happens when AI systems stop mainly serving humans and start negotiating, evaluating, and transacting with each other - and what content, media, and AI ops teams should build now.