AI/ML

I Handed Game Asset Production to an AI Agent. My Job Turned Out to Be Judgment

An AI coding agent produced six commercially usable, game-ready 3D props on my RTX 3090 in two rounds. It did the work, the measuring and most of the debugging. Four of the turning points came from me looking at something and saying 'no'. Here is where those happened and the review habits behind them.

·8 min read
#AI agents#Claude Code#human in the loop#image-to-3D#TRELLIS.2#Roblox#RTX 3090#workflow

TL;DR: I needed six small 3D props for a cat-themed Roblox obstacle course. I didn't model them, and I didn't write the pipeline. An AI coding agent (Claude Code, running on my RTX 3090) built a local image-to-3D pipeline, ran a little over 40 conversions, measured everything and delivered game-ready files.

My part was mostly short chat messages. But four of them changed the outcome more than anything the agent did on its own, and three of those rejected work that had passed every automated check. This post covers what I handed over, what the agent was better at than me, the four moments that needed a human, and the review habits I'd recommend to anyone directing an agent.

The technical side (pipeline, measurements, all seven bugs) is in the companion post: One Image to a Roblox-Ready 3D Prop on a Single RTX 3090.

The six final props

The setup

Two AI agents on two machines. A planning agent on another PC turns what I say into a written work order. A worker agent on the RTX 3090 machine picks the order up, does the work, and writes back a result file. I talk to both over chat and approve each work order before it goes out.

The work order for this job was short:

  • six props: kitten, rubber duck, life ring, paw-print raft, palm tree, beach parasol
  • only models whose licenses allow commercial use
  • one mesh per prop, at most 10,000 triangles, one 1024 px texture
  • the kitten lying down with its legs tucked in
  • previews from the front and side, triangle counts, and a log of every failure and retry

The most important lines in it came from things that had already happened to me.

What I brought that the agent couldn't

1. Knowing what I actually wanted

The first work order the planning agent wrote asked for promotional illustrations for the game. That was a reasonable reading of what I'd said, but not what I needed. I wanted 3D objects the game could use directly. When the illustrations came back, I said so, and the next order was for 3D.

An agent can execute a brief perfectly and still deliver the wrong thing. Checking the purpose, not just the output, is the first job.

2. Rules that came from real incidents

Two constraints in the brief came from experience:

  • One single mesh, legs tucked in. Earlier in the project, Roblox moderation had removed one of our AI-generated cat assets: a leg uploaded as a separate piece was flagged as sexual content. No model or agent would infer that rule. It exists because it happened to me.
  • Commercially usable models only. I'd already learned that an image model I'd been using for game art had a non-commercial license. So that rule came first, before any quality question. The agent then did the detailed checking, and found a catch I wouldn't have: the original TRELLIS pipeline bakes textures with an NVIDIA tool under a non-commercial license. It routed around that with ComfyUI's own reimplementation.

The person who carries the risk should set the non-negotiables. The agent is good at enforcing them.

3. "What is this?": rejecting on taste

The first kitten passed every check: one mesh, 8,931 triangles, the right size, the origin in the right place. I looked at it and laughed: "what is this?" It had a red cross where its mouth should be and blotchy shards across its body.

That rejection sent the agent digging. It found a remeshing mode that leaves a hidden inner shell under the surface, which pokes through once the mesh is simplified. One setting fixed the kitten and, as it turned out, several other models.

Even after that fix, I pushed again: "The quality is still too low. Is the model the problem?" The agent's answer was: partly the model, but more the pipeline. Try the free fixes before paying for another tool. Then it re-rendered the same mesh with proper lighting and shadows. It looked like a different asset. The mesh was fine. The flat preview renderer had been making it look cheap.

Same kitten, before and after the fix that my rejection triggered

Neither problem showed up in a number. They showed up because I said the result wasn't good enough and kept asking why.

4. "Is it turning its head?": spatial common sense

The fixed kitten (now 8,981 triangles, no shards) looked fine in its front preview. Then I noticed something: "The legs are pointing the wrong way. Is it turning its head?"

It was. The image generator had drawn the classic three-quarter pose: body lying sideways, head turned to the camera. That reads naturally in a photo, but as a 3D object it's a cat lying sideways with only its head turned toward you. The agent had reproduced the input faithfully, and every automated check passed.

The fixed kitten from five angles: the body lies sideways

I could see it because I know what a cat lying down looks like from every side. After that, the agent rendered every candidate from five angles and rejected any whose body went sideways.

What the agent was better at

I'd be misrepresenting this if I stopped there. The agent did things I couldn't have done in the time, or at all:

  • Volume and patience. About 50 input images and 40+ image-to-3D conversions at roughly 80 seconds each, with every run logged.
  • Root causes. For each rejection it found the mechanism behind it: the inner shell, faces folding over during simplification (it counted them: 1,099 folds at one setting, 11 at another), and a normal map that was 37% inside-out.
  • Measuring everything. Timing per stage and the pipeline's peak VRAM (18.5 GB, during image generation), and when it re-ran the same seed it found the 3D step isn't reproducible: 8,998 triangles one time, 10,029 the next.
  • Honest reporting. Every delivery listed the flaws that were still there: a tail a shade too red, a smudge on one parasol panel, a raft that looks more like plastic than wood. It didn't claim things were perfect.

The useful framing isn't human versus AI. The agent is fast and thorough at finding out why. I'm the one who knows whether it matters.

How I review an agent's visual work now

These came out of this project. They're what I'd tell anyone directing an agent on creative or visual tasks.

  1. Say the purpose, not just the deliverable. "3D props the game uses directly" is a different job from "pictures of props".
  2. Turn your incidents into rules. Anything that has burned you before goes into the brief as a hard constraint. The agent can't know your history.
  3. Ask for comparison sheets, not single results. Every candidate side by side, with the agent's pick marked and a one-line reason for each rejection. Reviewing a choice is faster and better than reviewing an output.
  4. Look from every side, in realistic lighting. One flattering preview hides problems. Five angles under game-like light expose them.
  5. Reject on taste, and say it in one sentence. "What is this?" was enough. You don't need to diagnose the problem yourself; the agent will. You do need to be willing to say no to something that passes the checks.
  6. Ask "is it the tool or the process?" before switching tools. I assumed the 3D model was the problem. What actually fixed it was a better preview.
  7. Sign off on the pick, don't just accept it. In the second round the agent picked the best kitten out of seven 3D candidates and marked it on the sheet. I agreed with it, but the sign-off was mine.

The result

Six props in two rounds, all one mesh, under 10,000 triangles, with one 1024 texture, made only with commercially usable models, on one consumer GPU. Four of them went straight into the game. The second round redid the raft and replaced the kitten with a clearer face, plus optional normal maps.

What I spent was attention, not modelling skill: a work order, a handful of chat messages, and a few honest reactions at the right time. I think that's the skill that matters as agents get more capable. The work is still there. What's left for me is deciding what good looks like and noticing when something isn't it.

The full technical write-up, with the pipeline, measurements, licenses and every bug, is in the companion post.

관련 글