Essential cookies keep Instavar working. Optional analytics help us understand how the site is used. Cookie Policy

Manage Cookie Preferences

Service reliability telemetry, including Sentry error monitoring and Vercel Speed Insights, stays enabled so we can secure the product and diagnose failures.

Skip to content
Instavar
ResearchPlaybooks
Research/Agents and coding models

Agents and coding models

Compare model behavior, agent reliability, tool use and the systems that evaluate them.

What we currently think

Model rankings depend on the work, constraints and evaluator. A useful comparison preserves failures and tests transfer beyond one prompt set.

Start with your question

  • Which model fits this kind of work?
  • Does the evaluation measure the property we care about?
  • What breaks when the agent operates a real system?

Start here

  • Kimi K3 vs GLM-5.2 vs GPT-5.6 Sol: Visual Benchmark

    Eight frontier-model routes explain one product, revise their work, and turn evidence into decisions. Compare the finished artifacts, costs, failures, and limits.

  • GPT-5.6 Luna vs Terra vs Sol: 294 Codex Runs Explained

    We compared GPT-5.6 Luna, Terra, and Sol across Codex reasoning levels using 294 ChatGPT-subscription runs. See why medium is the practical starting point, why no model won, and when higher reasoning may be worth the cost.

  • How We Used SkillOpt to Stress-Test Codex Skills

    We adapted the SkillOpt idea into a practical test loop for Codex skills: synthetic edge cases, narrow edits, and fresh unseen score streaks before we trusted an improvement.

Continue exploring

  • LFM2.5-2.6B Training Explained - SFT, Specialist Teachers, MOPD and Agentic RL

    How Liquid AI trained LFM2.5-2.6B for local agents using two-stage SFT, domain-specialist teachers, multi-domain on-policy distillation and reinforcement learning inside real agent harnesses.

  • Token Saver with Codex: A Smaller, Safer Way to Read Long PDFs

    We tested whether Codex could answer focused questions from long PDFs without receiving every page. See what worked, what did not, and where the approach fits in an AI video workflow.

  • Nanbeige4.2-3B vs Qwen3-VL-4B on an RTX 3090 Ti

    We ran 25 fixed arithmetic, planning, tool-call, and transformation tasks on Nanbeige4.2-3B and Qwen3-VL-4B. Here are the scores, latency, VRAM use, compatibility failure, and limits of this local agent benchmark.

  • Function Calling and MCP First Principles

    A plain-language walkthrough of how function calling and MCP fit together in a production agent loop.

  • Mining Claude Code and Codex Logs Into a Knowledge Base

    A practical guide to turning agent transcripts into reusable engineering memory without dumping every JSONL file into a vector database.

  • OpenAI Codex CLI /goal Command: What It Does and How to Enable It

    A source-inspected guide to the experimental OpenAI Codex CLI /goal command: what it stores, how it lets Codex continue work, how to enable the goals feature today, and why Codex Desktop users should treat it as CLI-only for now.

  • Running OpenAI Privacy Filter on an M2 MacBook Pro - 52-Case Benchmark

    Can OpenAI's new privacy-filter token classifier redact secrets and PII on a 16 GB M2 MacBook before they reach a cloud LLM? We ran 52 test cases covering API keys, PEM blocks, JWTs, names across six cultures, addresses, and decoy prose. Load time, latency, hit rates, and the AWS-key miss that matters.

  • Hardening Agents in Production - Locking Down the Attack Surface

    A live conference-note draft on why agents are already in production, why autonomous workflows create a larger attack surface than most teams expect, and what a real security control plane must do.

  • Systemic Decision Rot - Why AI Governance Must Move Beyond Model Safety

    A conference-note essay on why AI risk is not just about hallucinating models, but about how organizations transform AI outputs into summaries, recommendations, and institutional decisions without enough verification.

  • The Agent-to-Agent Internet - Evaluation Arenas, Algorithmic Governance, and the Dark Web of AI

    A strategy memo on what happens when AI systems stop mainly serving humans and start negotiating, evaluating, and transacting with each other - and what content, media, and AI ops teams should build now.

Singapore Office - Antiphishing Pte. Ltd.

JTC LaunchPad @ one-north

67 Ayer Rajah Crescent, #02-14

Singapore 139950

© 2026 Instavar. All rights reserved.

Find the angle that travels.

BlogResearchPlaybooksToolsCase StudiesAboutCareers|PrivacyTermsCookiesAI PolicyRights & ConsentReport Abuse

UK Corporate Presence - Litiga Ltd

Registered in England and Wales

Company number 11610573

Registered office: 128 City Road

London, England EC1V 2NX

Incorporated 8 October 2018