Review or Reviews
테크, 개발, AI, 하드웨어 — 실사용 기반 리뷰와 가이드
최신 글
I Handed Game Asset Production to an AI Agent. My Job Turned Out to Be Judgment
An AI coding agent produced six commercially usable, game-ready 3D props on my RTX 3090 in two rounds. It did the work, the measuring and most of the debugging. Four of the turning points came from me looking at something and saying 'no'. Here is where those happened and the review habits behind them.
One Image to a Roblox-Ready 3D Prop on a Single RTX 3090, Using Only Commercially Usable Tools
Six low-poly game props from FLUX schnell → TRELLIS.2 → Blender on one RTX 3090: about 80 seconds per model and an 18.5 GB VRAM peak. Seven things went wrong, from an inner shell that drew a red X on a kitten's mouth to a normal map that was 37% inside-out, and here is how each one was fixed.
I Fine-Tuned a 0.8B Classifier on Korean Business Plans. It Tied My Hand-Written Rules, Then Lost on New Data
A day on one RTX 3090 trying to make a 0.8B 'picker' model judge Korean startup grant applications. Synthetic training data scored 92/150; real text with targeted edits scored 126/150, tying the hand-written rules. Then a check I wrote down in advance failed on six new documents (48 vs 52).
더 보기
Qwen-Image-2.1 on One RTX 3090: 495 Generations Into Pixel-Art Game Sprites, Measured (and 6 Gotchas)
I moved every sprite for two small Android apps onto a local Qwen-Image-2.1 running on a single RTX 3090. 495 generations later: ~87 s per 1024² image, ~16 GB VRAM with CPU offload, and six gotchas that only showed up once the images had to work as real game assets.
I Built a Hallucination Gate for My Own Writing. The Bigger Local Model Didn't Catch More
A 40-claim eval for a local hallucination checker on 8B, 9B and 27B. The 27B caught fewer than the 9B, the union of the two small models beat every single model, and a frontier judge finished one claim ahead.
The Open-Model Cost Chart Everyone's Sharing Is API Prices. Here's What Self-Hosting Actually Gets You (Measured)
The intelligence-vs-cost chart making the rounds shows open models winning the value quadrant. True, but the x-axis is API token price. The cheap open winners are 100B-to-1T MoEs you can't run on a desktop GPU. Here's what you can actually self-host on an 11GB and a 24GB card, measured, and where the real ceiling is.
I Added a Verify Layer to My Local RAG to Catch Hallucinations. It Caught Me Being Wrong Twice About My Own Corpus
A claim-verification layer for a local RAG co-scientist, inspired by Karpathy's llm-wiki pattern. I tried to measure whether it catches hallucinations, almost shipped a false finding, and ended up with a clearer picture of what claim-checking can and can't do: it reliably catches values that are absent from the context, misses a real number pinned to the wrong question, and misses a false premise outright, and a model can't reliably referee its own blind spots.
What Actually Runs Well on a GTX 1080 Ti in 2026 (Measured)
The 'GPU poor' narrative says 24GB-and-below cards are eating well now thanks to QAT and MTP. But what about an 8-year-old 11GB GTX 1080 Ti? I measured it: Gemma 4 12B QAT at ~32 tok/s, Qwen3 8B at ~46, all fully on the GPU. Here's the table and where the ceiling is.
MTP Isn't Always a Win: 1.95× on My 3090, but Speculative Decoding Is Hardware-Dependent
MTP gave Gemma 4 12B QAT a 1.95x generation speedup on my 3090. But the same model with the same MTP draft runs 0.87x — slower — on an M1 Max. Speculative decoding is a hardware-dependent lever, not a free switch. Here are the measured numbers and why the draft-to-verify ratio decides it.