Community Discussion · Tracks

Jensen Huang Says AGI is Here: What Should You Test First?

Independent PanIndependent PanSep 72026/09/07 48 views

A friend recommended GPT-6 Astra to me, so I gave it a try. It's worth using, but don't trust the conclusions beforehand. Jensen Huang said it represents the arrival of AGI. You can treat it as a regular model upgrade and test it; decide whether to pay after testing. The term AGI is vague. In plain language, it's an AI that "can do a bit of everything." But your SaaS tools (software used monthly via browser), customer service emails, and data reports don't need it to do everything; they just need it to require fewer revisions for your specific tasks.

There is a typical segment in the source material.

“From ChatGPT to o1 to Astra in 4 years,” Huang said, “AGI has arrived.”

He also mentioned that Astra's training used over 100,000 Nvidia Grace Blackwell-related devices. These devices are typically the GPUs burned through during large model training. This information indicates "it's expensive, and they want to sell cards," not whether it can help you write correct customer emails. As someone who builds independent tools, I hate listening to launch events, so I'm in the habit of setting up a small test table first. Follow this process below; you can set it up in ten minutes and see if it's worth it in half an hour.

First, create a new local folder. Right-click on the desktop blank space, click New Folder, and name it agi-check. This folder is the drawer for your test materials. Once you see a yellow folder appear on the desktop, you're done.

Create prompts.txt inside the folder. Open it with Notepad and input 5 instructions you would actually use. Don't test "write a poem" or "explain quantum mechanics"; test your dirty work. For example, turning a customer complaint into a refund explanation, finding the error cause from a log snippet, translating a Chinese product page into English, explaining a code error, or writing 3 counter-arguments for a product page. A prompt is the instruction sent to the model; the closer it looks to your daily work, the more reference value it has. After saving, a text file will appear in the folder.

Next, create results.csv. CSV is a table separated by commas, which Excel, Numbers, and WPS can all open. Write the header task,model,output,time,fail,note in the first row, meaning Task, Model, Output, Time, Fail, Note. When you open the file, if you see column names separated by commas, you're good.

Open the GPT-6 Astra webpage, click New chat, and select GPT-6 Astra from the model dropdown. Copy the first item from prompts.txt into the input box and click Send. You'll see the answer appearing character by character. Wait for it to finish outputting, then copy the answer into the output column of results.csv. Here's a beginner pitfall: if the answer contains commas, the CSV will misalign. The fix is simple: wrap the answer in double quotes, or replace commas with vertical bars |. This table is quite handy, provided you don't mind its ugliness.

Record the time. From when you click Send until the last character appears, I usually use a phone stopwatch to estimate. Precision isn't necessary; obvious differences are enough. Then record failures: did it fabricate data, did the code fail to run, did it answer off-topic? Write yes for failure, no otherwise. I wrote a small tool myself to automatically record time when feeding prompts and answers, but for this tutorial, CSV is sufficient. After running 5 items, you'll see 5 rows of records in the CSV.

Switch to an old model or the free version and run the same set of prompts. This step is for control. Without a control, you might mistake "feeling good today" for "the model got stronger." In my tests, GPT-6 Astra indeed required two fewer follow-up questions for log summaries and error explanations, but when translating marketing copy, it still added drama, changing "simple and easy to use" to "revolutionary experience." This flaw exists in old models and hasn't been cured in new ones.

Let's discuss pitfalls separately. Don't just test pretty examples. Demo prompts are sanitized. What you usually input are requests with typos, half-sentences, and dirty data. Don't conclude based on one success. Run the same task at least 3 times. I first touched Rust 5 days ago; asking it to explain a memory error, the first time looked plausible, but the second time missed a rule. Don't immediately let an AI Agent automatically call your payment, email, or database APIs. An Agent is an AI that decides to call tools itself; an API is a channel for programs to call each other. Uncontrolled permissions aren't a joke; they can really turn a test into a production incident. Start with manual copy-pasting; it's safer.

Test your own messy work first, then listen to others talk about AGI. If the new model saves you one revision, one follow-up question, or one overtime session, it's worth using. If it just has more training cards and louder slogans, then for you, it's just material for vendor financial reports. After learning this, pick an AI feature you're about to pay for and use these 5 tasks as a pre-purchase test. If it fails, don't renew.


📌 This article is compiled from TechRadar. Original link: https://www.techradar.com/ai-platforms-assistants/chatgpt/jensen-huang-once-again-declares-that-agi-has-arrived-but-his-gpt-6-celebration-feels-like-hes-just-trying-to-sell-next-gen-nvidia-gpus

Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.

1 replies

?
Ctrl + Enter to reply
He Ma Chu Lai De

Let's not talk about AGI yet. If you can reduce fresh produce spoilage and improve restocking accuracy by another two percentage points, I'll believe you immediately.