I Fine-Tuned a 0.8B Classifier on Korean Business Plans. It Tied My Hand-Written Rules, Then Lost on New Data
A day on one RTX 3090 trying to make a 0.8B 'picker' model judge Korean startup grant applications. Synthetic training data scored 92/150; real text with targeted edits scored 126/150, tying the hand-written rules. Then a check I wrote down in advance failed on six new documents (48 vs 52).
TL;DR: I tried to replace part of a rule-based checker with Jeff, a 0.8B model that picks one of a few answers and gives a calibrated probability for each. The task is judging Korean startup grant applications question by question. On one RTX 3090:
- Synthetic training data (a 27B model writing fake plans) got 92/150.
- Real plan text with targeted edits got 126/150, tying my hand-written rules and well above the 108 you get by always answering "answered".
- I then wrote down a pass condition before scoring six never-seen plans: "rules + Jeff must beat rules alone". It failed, 48 vs 52.
LoRA training took ~25 minutes at 5.7 GB VRAM, and inference is ~0.1 s per question. The model isn't going into the app. The data recipe is the part worth keeping.
The task
Korean government startup programs (예비창업패키지 and friends) ask for a business plan in a fixed PSST shape: Problem, Solution, Scale-up, Team. I'm building a checker that reads a plan and answers 20 reviewer-style questions about it. Examples:
- "Is the market size backed by a sourced number?"
- "Is there a schedule table with a prototype date?"
- "Is there an unsupported 'first in Korea / no competitors' claim?"
Each question gets one of three verdicts, answered / weak / missing. Two "problem-finding" questions get present / absent.
The current engine is hand-written rules (pattern and structure checks) for 15 of the 20 questions, plus a language model for the other 5. The rules are fast and free, but brittle. My question was whether a small trained model could take over more of the work.
Why Jeff
Jeff is a 0.8B model built on Qwen3.5-0.8B. Its code is MIT and its weights are Apache 2.0. It doesn't generate text; it picks. You give it a state (here, the plan section), a question, and a set of labelled options. It returns a probability for each option, and a decision costs about as much as a forward pass. The repo ships a LoRA trainer and an "adapter kit" with split, leak and shortcut checks. That's exactly the shape of my problem.
One caveat up front: Jeff is English-first, and my documents are Korean. Everything below is me using it outside its comfort zone, not a judgment on what it was built for.
Setup
| GPU | RTX 3090 24 GB · i9-14900K · 192 GB RAM |
| Base | Jeff-Qwen3.5-0.8B v1.2 |
| Training | LoRA rank 16, alpha 32, lr 1e-4, 4 epochs, effective batch 16, max length 4,096 |
| Train time / VRAM | ~23–26 min per adapter, 5.7 GB peak |
| Inference | ~0.1 s per question (vs ~1.7 s for Qwen3.8 27B through Ollama on the same card) |
| Test set A | 10 official example plans (Korea Institute of Startup & Entrepreneurship Development, 2023, plus one consultant's public 2025 example) · 200 questions |
Two Windows notes if you try Jeff's trainer there. It assumes POSIX in four places: fcntl, fsync on a read-only handle, directory fsync, and replacing a symlink with os.replace. Each needs a one-line guard. Also, adapter names must be lowercase.
Round 1: synthetic data (92/150)
The obvious recipe: let a big local model (Qwen3.8 27B) invent 60 companies. For each plan section it first picks a target verdict per question, then writes text that should earn that verdict. Then the same model re-judges the text blind, and any disagreement is thrown away. About 60% survived, giving 715 rows. A second round added "PDF-ified" copies (no spaces between Korean words, wrapped lines, scattered table cells, masked numbers) to look like real extracted text.
Inside the synthetic world, the adapter hit 93–96%. On real plans it scored ~60%. It did learn something new, flagging "weak" answers at 20/36 where the base model managed 0. But on the full real test it reached 92/150, below the do-nothing baseline of 108.
The PDF-ification didn't help. The gap wasn't surface formatting. It was what the writer model thinks "weak" looks like versus what real plans actually contain.
Round 2: real text, edited (126/150)
So I flipped it: keep the text real and make the label come from an edit.

For each section of 15 real plans (the 10 above, plus 5 public plans found on GitHub), the first step is to mark the evidence lines for every question. Then:
- Delete a question's evidence lines → missing. Kept only if a re-read of what's left still finds no answer, which threw out 3–10 candidates per document.
- Rewrite them vaguely, with no numbers, sources or names → weak. Rejected if more than 2 digits survive.
- Paraphrase them, keeping every number and source → same label as before. This twin matters. Without it, the model can learn "AI-sounding sentence = weak".
- Insert one unsupported "first in Korea" sentence → present, for the problem-finding questions.
- Every other question whose evidence wasn't touched keeps its original label, so one edit yields several rows.
I started running these edits with Qwen3.8 27B too, then switched: as a judge it did no better than the untrained 0.8B (below), and the deletion check depends on judging well. The edit specs were instead written by five Claude sub-agents in Claude Code, each taking three documents, without access to the answer key. A script applied the edits mechanically.
Result: 989 rows:
| Kind | Rows |
|---|---|
| original text | 213 |
| paraphrase | 265 |
| deletion | 268 |
| vague rewrite | 210 |
| inserted claim | 33 |
To keep the test honest I cross-fitted. I split the 10 test plans into two halves. The adapter scored on half A was trained only on half B plus the GitHub plans, and vice versa, so every scored plan is unseen by the adapter that scored it.

92 → 126/150. Same model, same hyperparameters, same test; only the training data changed — invented text versus edited real text. That's the most useful finding of the day.
Two more things came out of the breakdown:
- Jeff held up across form changes. On the six plans written for a different program form (one my rules had never been tuned on), Jeff got 76/90. The rules got 67, and always-"answered" got 68. Rules quietly collapse when headings move; the model didn't.
- They were right in different places. Of the 150 questions, each one was right on 17 that the other missed. That looked like a combination should beat both.
Round 3: the check I wrote down first (48 vs 52)
Picking "which questions go to the rules, which to Jeff" by looking at that same test would be fitting the test. So before scoring anything new I committed this to the repo:
Problem-finding questions → rules. Everything else → Jeff (one adapter trained on all 989 rows). No confidence thresholds. Pass if the combination beats rules alone on new documents, and beats always-"answered".
New documents were hard to find; filled-in Korean plans online are mostly behind paywalls. I collected 9 more, mostly from GitHub. My app's section splitter failed on 3 of them: Roman-numeral headings, a table of contents mistaken for the body, and a document with no Solution heading. That's a real bug, and more important to users than anything in this post. That left 6 documents. Two independent Claude raters labelled them blind and agreed on 103 of 108 questions; the 5 disagreements were dropped.

The combination got 48/74, the rules alone 52/74, and Jeff alone 47. All three clear the 33/74 baseline, but the pass condition failed. A 4-question gap on 74 is within noise, but "within noise" is not "better", and better was the bar.

The per-label view shows why (the five label groups add up to the totals above: Jeff 47, rules 52). Jeff says "answered" too readily: when a section simply doesn't contain the thing, it said "missing" in only 4 of 16 cases, against 11 for the rules. On the problem-finding questions it caught all 4 real problems, but also "found" problems in 4 of 8 clean sections, flagging budget issues in sections with no budget at all. The rules got all 8 clean sections right.
So Jeff's one clear win, catching every "first in Korea" claim, came with false alarms on half the clean sections. Split the questions the way the pre-registration did and the rules come out ahead on both halves: 9 vs 8 of the 12 problem-finding questions, and 43 vs 39 of the other 62. Even with hindsight, flipping the split (Jeff on problem-finding, rules on the rest) gets 51, still short of the rules alone at 52.
What I'd tell myself yesterday morning
- Synthetic text doesn't transfer; synthetic edits of real text do. If you can't afford labels, edit real documents in ways whose effect on the label you control, and add a "same label" twin for every rewrite.
- Always print the do-nothing baseline. My first test set is ~72% "answered". Without the 108 line, Round 1's 92 would have looked like a win.
- Cross-fit, then pre-register, then test on new documents. The cross-fitted 126 was real but optimistic. The new documents told a smaller story, and only because the rule was fixed before I looked.
- A 27B local model judged no better than the untrained 0.8B (79 vs 78 of 150, at ~1.7 s per question). It marked well-written real answers "weak" far more often than the answer key did; bigger was not better on this task.
- Fix the boring parts first. The section-splitting bug costs more accuracy than any model swap.
Caveats
- The answer keys are not human-expert labels. The first test set was labelled by Claude in an earlier session, before any model results existed. The second was labelled by two independent Claude raters, and Round 2's edits were also written by Claude. A model sharing the labeller's taste may score a little high. A human consultant re-labelling a sample is the next step.
- Small n. 200 and 103 questions from 10 and 6 documents. Treat single-digit differences as noise.
- No document text is published here, only aggregates. Two of the public GitHub plans contain real names, so they're not identified.
- Jeff v1.3 (the announced long-term-support base) wasn't out yet. The recipe would carry over.
The app will keep the rules for the free, instant check and use a large hosted model for the paid deep review. The 0.8B adapter stays on the shelf, with a data recipe I'll reuse the next time a small model looks tempting.
관련 글
Best Ollama Models for RTX 3090 (2026): Qwen3 vs DeepSeek vs Llama Benchmarks
3월 30일 · 20 min read
Local LLMQwen-Image-2.1 on One RTX 3090: 495 Generations Into Pixel-Art Game Sprites, Measured (and 6 Gotchas)
10월 1일 · 7 min read
Local LLMMTP Isn't Always a Win: 1.95× on My 3090, but Speculative Decoding Is Hardware-Dependent
6월 11일 · 4 min read
Local LLMThe Prefill Wall: Why MTP's 2× Barely Moves Long-Context Latency (Qwen3.6-27B, RTX 3090)
6월 10일 · 6 min read