Research by Instavar
We test new AI tools on hardware people can actually buy.
These are our own experiments. Each article says what we tested, where we tested it, what happened, and where the result may not apply.
Written by Wei Jie Chee, founder of Instavar.
Latest experiments
- Original experiment
GPT-5.6 Luna vs Terra vs Sol: 294 Codex Runs Explained
We compared GPT-5.6 Luna, Terra, and Sol across Codex reasoning levels using 294 ChatGPT-subscription runs. See why medium is the practical starting point, why no model won, and when higher reasoning may be worth the cost.
coding agents · reasoning · model evaluation
- Original experiment16 GB M2 MacBook Pro and RTX 3090 Ti
Audio8 TTS 0.6B: Mac, RTX 3090 Ti and LoRA Fine-Tuning Test
We tested Audio8 TTS Preview 0.6B on a 16 GB M2 MacBook Pro and an RTX 3090 Ti, then fine-tuned it on 128 NSC-derived voice clips. See inference, memory, LoRA, ASR, speaker similarity, and checkpoint failure results.
text to speech · voice cloning · LoRA · consumer hardware
- Original experiment
LFM2.5 Encoder vs ModernBERT - CPU, GPU and Legal Document Benchmark
We compared LFM2.5-Encoder-230M and 350M with ModernBERT-base and large on CPU speed, RTX 3090 Ti latency, memory release, sentiment classification and long legal-document classification. Here are the results and their limits.
embeddings · document search · model evaluation · consumer hardware
- Original experiment
Nanbeige4.2-3B vs Qwen3-VL-4B on an RTX 3090 Ti
We ran 25 fixed arithmetic, planning, tool-call, and transformation tasks on Nanbeige4.2-3B and Qwen3-VL-4B. Here are the scores, latency, VRAM use, compatibility failure, and limits of this local agent benchmark.
vision language models · coding agents · model evaluation · consumer hardware
- Original experimentRTX 3090 Ti with 24 GB VRAM
MicroZoom on a 24 GB GPU - INT4, LoRA, and Image Quality Results
We tested MicroZoom inference and LoRA training on a 24 GB RTX 3090 Ti using INT4 weights. See how its images compared with Bicubic and Lanczos enlargement.
MicroZoom · image enlargement · LoRA · model quantization
- Original experiment
NuExtract3 Table OCR - How We Validate Numbers and Cell Structure
A public-safe FireRed and NuExtract3 comparison showing why table OCR needs numeric-token, row, column, and cell-position checks instead of visual review alone.
OCR · tables · document AI · model evaluation
- Original experiment
Surya OCR 2 Layout Analysis - Why Higher F1 Did Not Make It Our Default
Surya OCR 2 beat Paddle on layout F1 in two public-safe tests, but weaker label coverage, fragmented boxes, and fallback behavior kept it out of our default OCR pipeline.
OCR · document layout · document AI · model evaluation
- Original experiment
Running OpenAI Privacy Filter on an M2 MacBook Pro - 52-Case Benchmark
Can OpenAI's new privacy-filter token classifier redact secrets and PII on a 16 GB M2 MacBook before they reach a cloud LLM? We ran 52 test cases covering API keys, PEM blocks, JWTs, names across six cultures, addresses, and decoy prose. Load time, latency, hit rates, and the AWS-key miss that matters.
privacy · local AI · model evaluation · consumer hardware
- Original experiment
OmniDocBench Is Saturated - What Our 1,331-Page Benchmark Reveals About Real OCR Failures
OmniDocBench hits 94%+ accuracy across top models, but document parsing is far from solved. Our 1,331-page benchmark on scan-heavy PDFs exposes the failure modes that standard benchmarks miss: hallucinated text, broken table structures, spaced-letter artifacts, and blank-page blindness.
OCR · document AI · model evaluation
- Original experiment
How We Benchmark OCR Models on Scan-Heavy PDFs
A practical methodology for benchmarking OCR models on scan-heavy PDFs: corpus design, scoring, visual audit steps, failure modes, and the routing rule that emerged from a 31-PDF internal pilot.
OCR · scanned PDFs · document AI · model evaluation