Community Discussion · Company Watch

WorkBuddy Custom Model Integration: Image Input Failing, Anyone Else?

ShutterShutterSep 82026/09/08 87 views

I'm currently stuck trying to connect a vision-capable model to WorkBuddy. I've been using WorkBuddy for about a month, previously just using it to organize wedding photography and e-commerce product images by client and date. These past few days, I browsed the documentation on model configuration, saw automatic mode and custom APIs, and wanted it to help me classify images into 'dusk color temperature, cool tone, warm tone, portrait, product.' Result: Stuck.

I went to Settings > Model via my avatar, selected Custom API, filled in the URL and API Key. After saving, I could indeed see the model. But as soon as I ran product images from the asset library, it either gave only text descriptions or prompted that image input is not supported.

I tried toggling the custom protocol in advanced settings on and off, and things got weirder. Once it said the interface address was inaccessible, another time the task stopped halfway at 'Waiting for model response.' I also switched back to Automatic mode. It seemed to produce results, but the classification was scattered. It put a warm-light desk lamp into 'Dusk Color Temperature.' It looked plausible, but couldn't be delivered stably.

I'm not sure if the capability flags weren't checked correctly, or if the model name I entered was wrong. The docs say that when selecting standard vendors, capability flags like tool calling and image input are automatically written in, but I don't see obvious switches on my end. Maybe I misunderstood: WorkBuddy itself handles organization and delivery, while the model is just the foundation. But right now the foundation isn't connecting, and the entire asset library remains disorganized.

Has any veteran encountered this situation where saving succeeds but running images fails? Is it an API Key type issue, does the model not support multimodal, or should I not be doing recognition within WorkBuddy for this volume of e-commerce assets? Seeking guidance.

2 replies

?
Ctrl + Enter to reply
Long Ji
Long JiSep 8

I tried Qwen2-VL; WorkBuddy has extremely strict format validation for vision models. Switching to GLM-4V worked instantly.

Sister Qing

It's likely a Base64 format issue or the resolution is too high. When I hit this bug, switching to JPEG compression fixed it.