Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
A design brief as one image: can a 27B vision model build the animation? (automationoptimization.github.io)
3 points by airylizard 7 days ago | hide | past | favorite | 1 comment
 help



I made this after a fair complaint on an earlier post of mine: perplexity doesn't tell you whether a quantized model is usable. So the prompt here is a picture.

The model gets one PNG of a design brief and a one-sentence instruction to build it. Most of the brief is drawn rather than written. It has color swatches, a wireframe, dot counts shown as pips instead of numerals, a timeline strip in a deliberately unnatural order, and a few strings nobody could guess. If those strings show up in the output, the image was read. The output is a 12-second canvas animation in plain JS.

I ran six files of Qwen3.8-27B with the same cards, seeds and token budgets: bf16, llama.cpp Q4_K_M and Q3_K_M, my own 12 GB quant, and both Ternary Bonsai 2 files. Each reply is opened in headless Chromium with Playwright's fake clock stepped 40 ms per frame, so every file is graded at identical timestamps. It's also why the tiles play in sync. There are 17 checks per page, all fixed before any model ran. I hand-wrote a reference page per card plus mutants. TheI made this after a fair complaint on an earlier post of mine: perplexity doesn't tell you whether a quantized model is usable. So the prompt here is a picture. The model gets one PNG of a design brief and a one-sentence instruction to build it. Most of the brief is drawn rather than written. It has color swatches, a wireframe, dot counts shown as pips instead of numerals, a timeline strip in a deliberately unnatural order, and a few strings nobody could guess. If those strings show up in the output, the image was read. The output is a 12-second canvas animation in plain JS.

I ran six files of Qwen3.8-27B with the same cards, seeds and token budgets: bf16, llama.cpp Q4_K_M and Q3_K_M, my own 12 GB quant, and both Ternary Bonsai 2 files. Each reply is opened in headless Chromium with Playwright's fake clock stepped 40 ms per frame, so every file is graded at identical timestamps. It's also why the tiles play in sync. There are 17 checks per page, all fixed before any model ran. I hand-wrote a reference page per card plus mutants. The references score 17/17 and each mutant fails only the check it targets.

Mean page score out of 17, seven pages each, crashes counted:

  bf16    53.8 GB  13.9
  Q4_K_M  16.6 GB  13.7
  Q3_K_M  13.3 GB  13.0
  Mine    12.0 GB  11.4
  Bonsai   7.2 GB   5.9
  Bonsai   5.9 GB   5.0
On my own file: it's within 2 checks of Q4_K_M on 5 of 7 pages, and its lower mean comes from two pages on one card, one of which crashed on a block-scoping bug. A paired permutation test over the seven pages doesn't separate it from the other three Qwen files (p = 0.13 against Q4_K_M). With n = 7 that means "not resolved", not "equal". Bonsai's gap is resolved. 8 of its 14 pages threw JS errors, though the pages that ran had read the card correctly. I did not reproduce PrismML's own benchmark suite. The result I found most interesting isn't about quantization. No file, bf16 included, reads direction off a diagram. An arrow saying a bar grows right to left scored 0 of 8, and a counter-clockwise sweep given as numbered positions scored 1 of 8.

Caveats I know about:

The sample is tiny, so this is a demo with a rough score.

I changed the grader once after the first pass. Two files CSS-scaled a correct canvas to half size because the wireframe says "half scale", and the grader now captures the canvas at its drawn resolution. The page describes this.

I reopened every crashed page in a plain browser with no fake clock, and each one throws the same error there.

Sharing this partly because Bonsai 2 is getting a lot of attention for scoring 98% of BF16 on PrismML's own benchmarks.

On this task it didn't come close. references score 17/17 and each mutant fails only the check it targets.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: