TL;DR. I built a small, reproducible post-training experiment for an open-weight image model. A 16-image LoRA learned a cohesive emoji-like style for simple objects. It did not learn a new interaction such as a high-five. That boundary turned out to be the most useful result.
What we built
Emoji Studio now contains a deterministic dataset pipeline, a FLUX.2 Klein 4B LoRA recipe, retained checkpoints, blinded human checkpoint selection, a queued single-GPU inference service, transparent PNG/WebP export, and a private-alpha web interface.
The dataset used 16 training images and four validation images derived from Microsoft's MIT-licensed Fluent Emoji source. Every export has provenance, a checksum, and a caption. The source artwork was natively 256×256, so the experiment stayed intentionally bounded: upscaling cannot invent missing detail.
Training ran for 100 optimizer steps and saved adapters at steps 25, 50, 75, and 100. The successful replacement run took 108.61 seconds of optimizer time and used approximately 15,982 MiB of sampled peak GPU memory. Including recovery work, the quality-pilot stage cost approximately $0.41 by observed provider-credit change.
That is enough to demonstrate a real style shift. It is not enough to claim a universal design tool.
Later is not automatically better
The first evaluation compared base and adapted outputs on held-out objects, an unseen bicycle, and an excluded-domain raised-hand control. The adapted model moved toward softer, simplified plastic forms, but also removed useful detail. Both bicycle samples retained two wheels while regressing in crank, fork, or frame coherence.
We then rendered 48 images across checkpoints 25, 50, 75, and 100 on a concept-disjoint set. A blinded interface produced 72 pairwise human decisions. Step 25 won 24 comparisons, compared with 19 for step 50 and 10 each for steps 75 and 100. We therefore locked step 25 by checksum for serving instead of assuming the final checkpoint was best.
Serving exposed a different class of problems
On one 40 GB A100, the service completed all 18 jobs in its real GPU test: two deterministic smoke renders and 16 requests at concurrency four. It sustained 6.73 images per minute, measured 35.93-second p95 request latency, and used 16,983 MiB of peak sampled GPU memory. A warm generation took roughly eight to nine seconds.
The concurrency test exercised authentication, persistent jobs, polling, queueing, and backpressure around one serialized GPU worker. It did not make one diffusion pipeline generate four images simultaneously. That distinction matters when reporting throughput.
The service supports deterministic seeds, per-user hashed access tokens, owner-isolated artifacts, daily and alpha-wide quotas, feedback, immediate deletion, and seven-day retention. A managed serverless deployment was rejected because its compressed bundle exceeded a 5 MB code-store limit; the 32 MB LoRA was the dominant artifact. We kept the adapter private and used a bounded ordinary GPU instance with a localhost tunnel for manual testing.
The high-five failure was the clearest product lesson
When asked for hi5 emoji, the served model returned a polished generic smiley. It looked like an emoji and completely failed the task.
The shorthand prompt was not the real explanation. Our training set taught an object-oriented visual language. It did not contain examples needed to learn two hands meeting palm-to-palm, contact geometry, finger anatomy, or the semantic distinction between waving and high-fiving.
A style LoRA can change how a known concept looks. It does not automatically add a missing compositional skill. Seeds sample other points in the existing distribution; they do not create a capability that was never represented in training.
What it cost
Before the live private-alpha session, the whole-project estimate was approximately $3.83, including baseline inference experiments, training preparation and recovery, the quality LoRA, checkpoint selection, and the service load test. Provider accounting can lag, so these are observed credit deltas rather than finalized invoices.
The surprisingly expensive parts were model downloads, failed hosts, uploads, evaluation rendering, artifact recovery, and infrastructure validation—not the 108 seconds of successful gradient updates.
What comes next
The next experiment focuses narrowly on high-fives. We will freeze an evaluation set covering hand count, finger readability, palm orientation, contact point, skin-tone variation, small-size readability, and absence of extra limbs. We will measure the base model and prompt-only ceiling before curating or commissioning licensed interaction examples.
For memes, the image model should generate the illustration while HTML, Canvas, or another normal graphics layer handles typography. That gives us exact spelling, accessible text, editable layouts, and deterministic exports.
The strongest claim we can make today is simple: we built and exercised a reproducible post-training system, measured where it worked, preserved where it failed, and used those failures to define the next experiment.