Bringing Google's compute back to California? I tried running models locally with GGUF
Moving Google's compute back to California? I tried running models locally with GGUF
The day before yesterday (August 2nd), I saw a news piece saying that Google moved a huge chunk of AI compute from other locations to California to compete with Anthropic and OpenAI for talent and resources. Basically, talent is flocking to the Bay Area, and servers are following suit westward. It got me thinking: what's the underlying logic behind these big tech battles? Why is compute so critical?
As a developer, although I wrote a llama.cpp tutorial before, it was mostly about setting up the environment. This time, I tried running models locally using the GGUF format, helping beginners truly get "up and running," while also understanding the strategic significance of where compute is located.
Day One (Aug 2): Understanding What Compute Is
First, you need to know that AI models are essentially piles of math formulas, and running them requires computational resources. Training models needs thousands of GPUs queued up to calculate, typically using professional cards like H100s or A100s. But if you're just using the model (this is called inference), you can run it on a regular computer—it's just slower.
The tools I used are llama.cpp + GGUF format. llama.cpp is a program written in C++ designed specifically to run models on ordinary computers. GGUF is a model format, similar to a zip archive but optimized specifically for llama.cpp.
Step 1: Download llama.cpp. Open your browser, search for "llama.cpp releases," find the latest version, and download the file corresponding to your system (for Windows, choose llama.cpp-xxx-win64.zip). Extract it to your desktop and create a folder named llama.
Step 2: Open PowerShell and cd into that folder (right-click on empty space in the folder and select "Open in Terminal").
Step 3: Type ./llama-cli.exe to see if it throws an error. If you get "VCRUNTIME140.dll not found," go to Microsoft's official site, search for "Visual C++ Redistributable," and install it. I got stuck here on day one and almost smashed my computer out of frustration.
Day Two (Aug 3): Downloading Models and Getting Them Running
Where do you find model files? Go to HuggingFace and search for "GGUF" plus the model name you want. I recommend beginners start with a GGUF file for a small model (e.g., Q4_K_M quantization); it's under 1GB and runs fast.
On HuggingFace, click the filename, then click the "Download" button. If the download fails, you can use the huggingface-cli command-line tool: huggingface-cli download TheBloke/a-small-model-GGUF-format a-small-model-GGUF-file-e.g-Q4_K_M --local-dir ./models. Once the progress bar finishes, move it into the models subfolder inside your llama folder.
Then get it running: In PowerShell, type ./llama-cli.exe -m ./models/a-small-model-GGUF-file-e.g-Q4_K_M -p "Please explain what compute power is in one sentence". You'll see a bunch of logs, and then the model will answer you. When I first saw that response, it felt like I had built my own mini data center.
Pitfall: If your GPU is NVIDIA, remember to add the -ngl 35 parameter, which means offloading 35 layers to the GPU for acceleration. Without it, it runs on the CPU and is painfully slow. I didn't figure this out until day two; before that, I thought my computer was just too weak.
Today (Aug 5): Understanding Why Compute Matters
Now you understand why Google is moving compute to California. Because the larger the model, the more compute it requires. Yesterday (Aug 4), I tried a 70B model, and it took several minutes to process a single sentence—let alone training. The competition between big companies is fundamentally about who can stack more compute faster.
I tested for three days, scaling from 1B models to 7B models, and noticed a pattern: the bigger the model, the better the performance, but memory and VRAM requirements grow exponentially. My lab's 3090 can only handle 7B models; anything larger causes VRAM overflow.
So, Google moving compute to California is superficially about talent flow, but actually involves the physical migration of infrastructure. Where the compute is, the talent is there, and technical iteration happens there.
After learning this, the next step could be trying to adjust thread counts using the -t parameter in llama.cpp to see how to configure hybrid CPU/GPU inference. Or try running a 7B model to experience what "quantitative change leads to qualitative change" feels like.
📌 This article is compiled from Bloomberg Tech. Original: https://www.bloomberg.com/news/articles/2026-08-06/google-shifts-ai-power-to-california-in-race-against-anthropic-openai
Copyright belongs to the original authors. This is a compilation and independent analysis based on public reports.
Physix Frontier