After a Month with WorkBuddy: It Saved Me Time on AI Video Model Benchmarks, But Not My Taste
I hate "model comparison reviews." It's not that I dislike the models; it's that I hate seeing a pile of parameters, screenshots, group chats, and PDFs stacking up on my desktop like a stale latte. In early September, I need to evaluate internal video generation tools. The requirements are typical: Ops wants to see which open-source model can run locally, Legal only cares about the license, Dev looks at VRAM, and I just want to know if the output is watchable. I scrolled past an AI video model comparison review where five locally deployable models were listed densely—resolution, duration, audio, parameters, architecture—it looked like a medical checkup report at first glance. My initial reaction wasn't "how comprehensive," but rather, "if I had to piece this together myself, I'd be working myself to death." So I threw that review, several model cards, screen recordings from colleagues, and email constraints into WorkBuddy, hoping it would do the dirty work first.
For the past month, I've mainly used WorkBuddy for document archiving, table merging, email summaries, and meeting minutes. A few days ago, I wrote about how its mixed input messed up field mapping, and I stepped on that rake again. At first, it was very enthusiastic, like a bug that just woke up, sweeping everything off the desk onto the floor. LTX 2.5's 4K got shoved next to MiniMax H3, HappyHorse's API attributes were labeled as locally deployable, and the license column just vanished. The most absurd part was that it standardized "max duration" to just "duration," making 5 seconds and 15 seconds look like the same category. At that moment, I almost closed the software. I also nitpick the UI; the default font size in exported tables looks like a stale latte, with gray backgrounds and white text blurring into one mess. But ugly as it is, it did break down twenty-plus files into rows.
Later, I learned my lesson and stopped expecting it to generate a "correct comparison" in one step. I gave it a field contract, telling it not to rush to conclusions but to create an evidence table first: model name, source, original quote, extracted fields, confidence level, items to confirm. This operation is very WorkBuddy-like; it doesn't give direct answers like model benchmarks but can move scattered things into the same basket. I told it to process each document individually, forbidding cross-source merging; once it separated the sources cleanly, then let it map the fields. The accuracy improved significantly on the second pass. My tests showed about 70% usability on the first pass, nearly 90% on the second. The remaining 10% were all semantic traps: open-source weights, locally deployable, API available, commercial license, default resolution, max resolution—these terms often impersonate each other in documentation.
This shows WorkBuddy's position. It isn't doing the selection for me; it's sorting out the garbage before the selection. AI video model comparisons look like technical evaluations but are actually information governance. The prettier the parameter table, the easier it is to deceive. One model says "supports 4K," another says "max resolution 4K (requires specific config)," and a third says "community version can run locally, official API is more stable." Putting these sentences in the same table without sources is just being shady. WorkBuddy's benefit is that it keeps the source column; its drawback is that it has no aesthetic sense or responsibility—it won't judge which statement is trustworthy for you. WorkBuddy's real value isn't doing the selection for me, but sorting out the pre-selection garbage. I wanted to say this when I wrote that it "shouldn't compete on beauty with three-second posters," and this time I'm even more sure.
When I exported the organized field table into a PPT outline, WorkBuddy gave me three pages: capability matrix, risk list, pilot suggestions. The content was usable, but the layout was ugly. Title spacing looked misaligned, table borders looked drawn with a ruler by hand. I deleted the color scheme after two seconds. But I admit, what it saved me wasn't ten minutes of "making a table," but forty minutes of "finding evidence." Previously, this kind of work was the most fragmented: a table in a review article, a paragraph in a README, a line in a group chat saying "this runs," and an email adding "Legal says the license doesn't work." After WorkBuddy took over this dirty work, I could spend time on what truly needs humans: deciding which fields to highlight, which models matter to the team, and which charts shouldn't mislead for the sake of looking good.
This also reminds me of recent video model leaderboards. Few in the Top 10 can run locally, and the open-source track competes on parameters, VRAM, duration, audio, and licenses, with every dimension being fiercely contested. But office tools compete differently. WorkBuddy doesn't generate 4K videos, nor does it train models for you; it competes on the messy inputs on your desktop. If a model parameter is wrong, it might just mean the benchmark isn't rigorous; if a field mapping is wrong, it directly becomes a promise in meeting minutes, a conclusion in a PPT, or blame-shifting in emails. So now, when I look at AI office tools, I care less about whether they can "generate stunning reports" and more about whether they can block errors. For example, field contracts, source traceability, rollback permissions. I even think hardcoding rollback permissions into the workflow is more effective than writing documentation. When WorkBuddy made mistakes on the first pass, I could revert to the previous version of the evidence table instead of digging through the file pile again, which was more useful than giving me a pretty sentence.
However, WorkBuddy has unsuitable areas too. It's not suitable for directly producing the final selection report. It can organize materials into structured drafts but won't take on judgment for you. Especially regarding licenses, local deployment, and commercial restrictions, it occasionally mixes similar concepts, like blending two grays into one dirty color. It's also not suitable for purely visual assets; ask it to layout a comparison chart, and it will give you a result that is logically correct but aesthetically disastrous. What UI designers fear most isn't ugliness, but ugliness that makes sense—where every cell in the table is correct, yet the whole thing looks like it never went through design review. WorkBuddy's interface and functionality have a bit of this vibe: clear functional paths, rough visual details, button spacing looking like someone dragged them, and lazy export templates. I dislike it, but I can't live without it at work. That's probably the relationship between adults and tools.
Looking ahead, I feel these comparisons will force tools like WorkBuddy to become more "bounded." Currently, AI video models compete on parameters, while office tools compete on workflows. Whoever can break multi-source inputs into evidence chains, whoever can make field mappings verifiable, will transform from a "generator" to an "executor." I don't hope WorkBuddy becomes an AI that writes prettier conclusions; I hope it becomes more like a meticulous archivist. Every field must have a source, every merge must have conflict warnings, every export must keep the previous version, and every error must be reversible. The interface definitely needs changes too—stop using those gray-background blurred-text templates. At least clean up table alignment, font hierarchy, and source column width. I can't use something ugly; that's the truth.
If you're also using WorkBuddy for similar comparisons, selections, or multi-document merges, my advice is straightforward: Don't treat it as a report generator; treat it as a sorter. First, give it a field contract, then let it create an evidence table, and finally manually verify key fields. Separate "local deployment" from "open-source weights," separate "max resolution" from "default output," and separate "license" from "scope of permission." Make every row carry a source, and tell it not to summarize yet. You can accept its first-pass chaos, but give it a correction path for the second pass. Tools lock down the floor; humans fill in the imperfections. Bug just jumped on the keyboard and deleted my exported table. I'm too lazy to scold him; perfect timing to rerun the rollback.
Physix Frontier