Essential cookies keep Instavar working. Optional analytics help us understand how the site is used. Cookie Policy

Manage Cookie Preferences

Service reliability telemetry, including Sentry error monitoring and Vercel Speed Insights, stays enabled so we can secure the product and diagnose failures.

Skip to content
BlogTools
  1. Home
  2. Blog
  3. Research

Singapore Office - Antiphishing Pte. Ltd.

JTC LaunchPad @ one-north

67 Ayer Rajah Crescent, #02-14

Singapore 139950

© 2026 Instavar. All rights reserved.

Find the angle that travels.

BlogResearchToolsCase StudiesAboutCareers|PrivacyTermsCookiesAI PolicyRights & ConsentReport Abuse

UK Corporate Presence - Litiga Ltd

Registered in England and Wales

Company number 11610573

Registered office: 128 City Road

London, England EC1V 2NX

Incorporated 8 October 2018

Research by Instavar

We test new AI tools on hardware people can actually buy.

These are our own experiments. Each article says what we tested, where we tested it, what happened, and where the result may not apply.

Written by Wei Jie Chee, founder of Instavar.

Latest experiments

  1. Original experiment

    GPT-5.6 Luna vs Terra vs Sol: 294 Codex Runs Explained

    We compared GPT-5.6 Luna, Terra, and Sol across Codex reasoning levels using 294 ChatGPT-subscription runs. See why medium is the practical starting point, why no model won, and when higher reasoning may be worth the cost.

    coding agents · reasoning · model evaluation

  2. Original experiment16 GB M2 MacBook Pro and RTX 3090 Ti

    Audio8 TTS 0.6B: Mac, RTX 3090 Ti and LoRA Fine-Tuning Test

    We tested Audio8 TTS Preview 0.6B on a 16 GB M2 MacBook Pro and an RTX 3090 Ti, then fine-tuned it on 128 NSC-derived voice clips. See inference, memory, LoRA, ASR, speaker similarity, and checkpoint failure results.

    text to speech · voice cloning · LoRA · consumer hardware

  3. Original experiment

    LFM2.5 Encoder vs ModernBERT - CPU, GPU and Legal Document Benchmark

    We compared LFM2.5-Encoder-230M and 350M with ModernBERT-base and large on CPU speed, RTX 3090 Ti latency, memory release, sentiment classification and long legal-document classification. Here are the results and their limits.

    embeddings · document search · model evaluation · consumer hardware

  4. Original experiment

    Nanbeige4.2-3B vs Qwen3-VL-4B on an RTX 3090 Ti

    We ran 25 fixed arithmetic, planning, tool-call, and transformation tasks on Nanbeige4.2-3B and Qwen3-VL-4B. Here are the scores, latency, VRAM use, compatibility failure, and limits of this local agent benchmark.

    vision language models · coding agents · model evaluation · consumer hardware

  5. Original experimentRTX 3090 Ti with 24 GB VRAM

    MicroZoom on a 24 GB GPU - INT4, LoRA, and Image Quality Results

    We tested MicroZoom inference and LoRA training on a 24 GB RTX 3090 Ti using INT4 weights. See how its images compared with Bicubic and Lanczos enlargement.

    MicroZoom · image enlargement · LoRA · model quantization

  6. Original experiment

    NuExtract3 Table OCR - How We Validate Numbers and Cell Structure

    A public-safe FireRed and NuExtract3 comparison showing why table OCR needs numeric-token, row, column, and cell-position checks instead of visual review alone.

    OCR · tables · document AI · model evaluation

  7. Original experiment

    Surya OCR 2 Layout Analysis - Why Higher F1 Did Not Make It Our Default

    Surya OCR 2 beat Paddle on layout F1 in two public-safe tests, but weaker label coverage, fragmented boxes, and fallback behavior kept it out of our default OCR pipeline.

    OCR · document layout · document AI · model evaluation

  8. Original experiment

    Running OpenAI Privacy Filter on an M2 MacBook Pro - 52-Case Benchmark

    Can OpenAI's new privacy-filter token classifier redact secrets and PII on a 16 GB M2 MacBook before they reach a cloud LLM? We ran 52 test cases covering API keys, PEM blocks, JWTs, names across six cultures, addresses, and decoy prose. Load time, latency, hit rates, and the AWS-key miss that matters.

    privacy · local AI · model evaluation · consumer hardware

  9. Original experiment

    OmniDocBench Is Saturated - What Our 1,331-Page Benchmark Reveals About Real OCR Failures

    OmniDocBench hits 94%+ accuracy across top models, but document parsing is far from solved. Our 1,331-page benchmark on scan-heavy PDFs exposes the failure modes that standard benchmarks miss: hallucinated text, broken table structures, spaced-letter artifacts, and blank-page blindness.

    OCR · document AI · model evaluation

  10. Original experiment

    How We Benchmark OCR Models on Scan-Heavy PDFs

    A practical methodology for benchmarking OCR models on scan-heavy PDFs: corpus design, scoring, visual audit steps, failure modes, and the routing rule that emerged from a 31-PDF internal pilot.

    OCR · scanned PDFs · document AI · model evaluation