Community Discussion · Tracks

Running tumor AI on cloud GPUs: A week of pitfalls

YimingYimingSep 82026/09/08 89 views

Last week, our team took on an internal validation project to see if AI could assist in filtering tumor-immune-related data. The pain point was direct: opening a training script on a local M2 Max caused the fans to roar, and before the data finished running, VRAM usage spiked red. As a 30-person startup, we're terrified of buying a pile of GPUs for a small demo, so I built a workflow using cloud GPUs I've been using for about a month.

On Day 1, I focused on data cleaning. I used GLM Coding Plan, which I've been trying for less than a week. It generates Python scripts quickly; aligning fields like sample IDs, mutation sites, and expression levels produced basic code in minutes. But bottlenecks appeared fast. Public medical data has missing values, inconsistent units, and messy field names. The generated code looked correct but crashed with index errors upon execution. I asked it to fix it; after two or three rounds, I ended up adding logs and exception branches myself. However, dumping the error stack back to it allowed it to quickly locate whether issues were null values or field mapping errors, saving junior engineers from repeatedly checking documentation.

On Day 3, I threw the cleaned small-sample dataset onto cloud GPUs. Cloud GPU means computing models on remote server graphics cards, avoiding self-built server rooms. I've been testing Lambda recently, a cloud GPU rental service. In the interface, task status moved from queued to running, but default batch sizes blew out VRAM immediately. Only after shrinking the batch did it run through. What really made me pause was queuing and quotas. I recently read Arm CEO Rene Haas saying AI cancer therapies are slowed by chip shortages, noting that modeling how DNA markers (gene variants or expression features) are affected by cancer is currently too complex. After running a cycle, I viewed this through a startup's ledger lens. You're calculating HBM (high-bandwidth memory), storage chips, supply chains, power consumption, and reservable compute windows. I just started touching domestic GPUs yesterday; adaptation costs aren't fully calculated, so I wouldn't dare rely on them for firefighting.

A week later, I broke this validation down into a small pipeline. Not chasing volume, but ensuring the flow from cleaning, training, inference to manual review works. I also tried an AI Agent I'd just picked up two days prior—a model program capable of automatically running tasks—to organize experiment logs and failure reasons, but I didn't let it automatically modify models or submit results. Beyond automation, manual review and log auditing must be retained. When terms like T-cells and pMHC first appeared, someone on the team asked, so I explained them as immune cells identifying abnormal cells via surface fragments, with AI helping humans narrow candidate ranges faster. An Atlantic Monthly article argued AI is unlikely to cure cancer soon; I found this reminder useful. AI accelerates screening, but therapeutic judgment still requires humans.

Conclusion: It depends. If you already have a data team capable of handling dirty data and clinical/bioinformatics partners, and just want prototype validation, this cloud GPU + AI-assisted coding setup is worth trying. It compresses weeks of environment setup and script debugging into days. If you expect it to directly discover cancer therapies, or if your cash flow is too tight to bear hidden costs like queuing, adaptation, and manual review, I don't recommend it. Implementation requires building a small pipeline from cleaning, training, inference to manual review first. Competition in medical AI may shift from models to data pipelines and compute scheduling. Stable reproduction of results in small samples matters more than stacking models.


📌 This article is compiled from Hacker News. Original link: https://www.bbc.com/news/articles/c0m39g7xzevo

Copyright belongs to the original authors. This text is a compilation and independent analysis based on public reports.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts