Review or Reviews

테크, 개발, AI, 하드웨어 — 실사용 기반 리뷰와 가이드

최신 글

더 보기

I Fine-Tuned a 0.8B Classifier on Korean Business Plans. It Tied My Hand-Written Rules, Then Lost on New Data

A day on one RTX 3090 trying to make a 0.8B 'picker' model judge Korean startup grant applications. Synthetic training data scored 92/150; real text with targeted edits scored 126/150, tying the hand-written rules. Then a check I wrote down in advance failed on six new documents (48 vs 52).

10/4

Qwen-Image-2.1 on One RTX 3090: 495 Generations Into Pixel-Art Game Sprites, Measured (and 6 Gotchas)

I moved every sprite for two small Android apps onto a local Qwen-Image-2.1 running on a single RTX 3090. 495 generations later: ~87 s per 1024² image, ~16 GB VRAM with CPU offload, and six gotchas that only showed up once the images had to work as real game assets.

10/1

I Built a Hallucination Gate for My Own Writing. The Bigger Local Model Didn't Catch More

A 40-claim eval for a local hallucination checker on 8B, 9B and 27B. The 27B caught fewer than the 9B, the union of the two small models beat every single model, and a frontier judge finished one claim ahead.

8/10

The Open-Model Cost Chart Everyone's Sharing Is API Prices. Here's What Self-Hosting Actually Gets You (Measured)

The intelligence-vs-cost chart making the rounds shows open models winning the value quadrant. True, but the x-axis is API token price. The cheap open winners are 100B-to-1T MoEs you can't run on a desktop GPU. Here's what you can actually self-host on an 11GB and a 24GB card, measured, and where the real ceiling is.

6/23

I Added a Verify Layer to My Local RAG to Catch Hallucinations. It Caught Me Being Wrong Twice About My Own Corpus

A claim-verification layer for a local RAG co-scientist, inspired by Karpathy's llm-wiki pattern. I tried to measure whether it catches hallucinations, almost shipped a false finding, and ended up with a clearer picture of what claim-checking can and can't do: it reliably catches values that are absent from the context, misses a real number pinned to the wrong question, and misses a false premise outright, and a model can't reliably referee its own blind spots.

6/19

What Actually Runs Well on a GTX 1080 Ti in 2026 (Measured)

The 'GPU poor' narrative says 24GB-and-below cards are eating well now thanks to QAT and MTP. But what about an 8-year-old 11GB GTX 1080 Ti? I measured it: Gemma 4 12B QAT at ~32 tok/s, Qwen3 8B at ~46, all fully on the GPU. Here's the table and where the ceiling is.

6/12