🖼️ Qwen3.5 VLM caption evaluator

Drop images, pick a model, compare captions. One model loads at a time (~1 min on first use / when switching).

Model
32 1024