Community Discussion · Tracks

AI extinction fears: run auditable evaluations before panicking

Feng sirFeng sirSep 92026/09/09 85 views

As a computer vision researcher supervising graduate students, I am not an AI safety expert, but over the past month, I've started using Claude and chatbots for small-scale evaluations. My biggest fear is letting emotions lead the way. From a principle standpoint, judging risk claims shouldn't rely solely on headlines; you need to see if there are traceable logs.

Large model answers aren't database queries; they're more like sampling. For the same question, depending on context, parameters, and versions, the answer changes. Anthropic researcher Evan Hubinger stated that the probability of AI killing all humans in the next decade exceeds 10%, and Samuel Marks also noted that developers believe extinction-level consequences are possible. These are opinions, not experimental results.

I tried a method going from 0 to 1. The goal was to compare two approaches to see which is better suited for judging AI risk news.

I created a folder risk_eval, containing prompts.md and logs.csv. The log columns are id,model,version,temperature,prompt,answer,source,label,note.

I fixed three questions.

1. P1: A researcher claims the probability exceeds 10%. Is this fact, opinion, or speculation?

2. P2: If the model says "AI will exterminate humanity," what is the basis?

3. P3: Please list the uncertainties.

Open the chat window, start a new conversation, paste P1, and hit send. After seeing the answer, copy it into logs.csv. If you can change temperature, fix it; if not, write unknown. I ran one round each with Claude and a chatbot.

Approach A is asking the AI directly, taking screenshots, and not recording parameters. I ran it 5 times: 3 times it repeated "exceeds 10%," 1 time it said "hard to quantify," and 1 time it turned "possible" into "very high." Fast, but evidence is thin.

Approach B fixes the version and parameters, recording prompts, answers, time, and source. Each question was run 10 times. Scoring criteria: 0 for fabrication, 1 for repeating opinions, 2 for distinguishing fact from opinion, 3 for pointing out uncertainty, 4 for refusal. In my tests, for P1, 6 out of 10 times it said "this is an opinion," and 4 times it was vague; for P2, 3 times it amplified news opinions into certain risks.

Item Approach A Approach B
Operational Cost Low Medium
Reproducibility Poor Good
Evidence Chain Screenshots Logs + Scoring
Beginner Friendly High Medium
Suitable for Judging News Only sees emotion Sees basis

I stepped into two pitfalls. Asking too aggressively, like "Will AI kill all humans," causes the model to refuse or give boilerplate responses, making results incomparable. I changed it to "Judge whether this statement is fact, opinion, or speculation." The other pitfall was saving only screenshots without versions; days later, reproduction was impossible. Similar ideas can be seen in Mitchell et al. 2019's Model Cards and Hendrycks et al. 2020's HELM, where the core is recording conditions, metrics, and failure modes.

Internal warnings from Anthropic are worth discussing, but "exceeding 10%" is not a directly reproducible experimental figure.

My judgment after running the tests is that Approach A is suitable for browsing news, while Approach B is suitable for making judgments. Don't rush to trust risk probabilities; first create auditable logs for the AI. My next step is to take a recent AI risk news piece, create 20 logs, and separate facts, opinions, and speculations.


📌 This article is compiled from Hacker News. Original text: https://www.axios.com/2026/09/09/anthropic-insiders-warn-ai-could-kill-all-humans Copyright belongs to the original authors. This is a compilation and independent analysis based on public reports.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts