As a Chinese Lit sophomore, I let AI analyze 96 posts to finalize my topic selection that night
Community Discussion · Policy

As a Chinese Lit sophomore, I let AI analyze 96 posts to finalize my topic selection that night

XieXieSep 92026/09/09 133 views
IMG_256

As a sophomore in Chinese Language and Literature, I had AI read through 96 posts for me, and my thesis topic was set that same night

My graduation thesis last semester nearly went off the rails.

I spent two weeks going back and forth with my advisor on the topic, downloaded 37 papers but only read the abstracts of each, struggled for a week to write less than 3,000 words for the first draft, got a plagiarism check score of 42%, revised until 4 AM, and after finishing, I couldn't even make sense of it myself. With three days left until the deadline, I hadn't even finalized the introduction.

Then my roommate threw a WorkBuddy at me. I used it as a Hail Mary pass for three days, and ended up scoring 85 on my thesis. My advisor's comment was "clear structure, sufficient argumentation."

It wasn't until then that I realized: writing a thesis isn't about you not working hard enough; it's that your tools don't match your effort. Others are already using AI to build frameworks, write drafts, reduce plagiarism rates, and fix formatting, while you're still crafting every character by hand.

Today I'm laying out my complete workflow from those three days. I'll be honest about the pitfalls I hit and the mistakes I made.

1. Don't rush to flip through literature; let AI capture "what everyone is talking about" first

A quick note on the tool. WorkBuddy understands plain Chinese. You just specify what to grab, how to categorize it, and what tags to apply, and it scrapes discussions from social platforms and organizes them into a data table you can view directly. No coding required—even a pure humanities student like me dared to try it.

Where I got stuck was not knowing what real readers were actually discussing regarding Dream of the Red Chamber. CNKI (China National Knowledge Infrastructure) is full of scholars' discussions, but you can't see what living people are talking about. Sending out surveys is too slow, and interviews aren't feasible due to lack of participants. So I gave it this prompt:

Search Xiaohongshu (Little Red Book) for posts from the past week, keywords including: Dream of the Red Chamber, Lin Daiyu, Xue Baochai, Dream of the Red Chamber character discussions. Scrape post content and comments, categorize by four dimensions: character discussion, plot controversy, film/TV adaptations, and interpersonal dynamics. For each entry, extract title, body summary, comment count, like count, publish time, and IP location, add discussion tags, and finally tally counts by category.

Then I went to get a cup of water. When I came back, 96 entries were listed on the screen: 32 character discussions, 24 plot controversies, 24 interpersonal dynamics, and 16 film/TV adaptations. Each included title, summary, comment count, like count, publish time, and IP location. The AI also added tags: shipping CPs, workplace etiquette, art of conversation, adaptation complaints, reading notes.

One sentence, three minutes, 96 entries. You can copy this step directly: just swap the keywords for topics relevant to your major. For sociology, use social issues; for education, use learning confusions; for journalism, use hot events. Keep everything else the same.

2. The data overturned all the topics I thought I knew

Don't rush to use the data once you have it. The first thing I did was write down my original three assumptions, then test them against the data one by one. All three were shattered.

I assumed the hottest discussion would definitely be the love story between Baoyu and Daiyu. There were indeed many posts related to Daiyu, but when I isolated the ones with high save counts, what was repeatedly saved and cited by office workers was "Xue Baochai's way of handling things." Working professionals were treating it as a textbook on interpersonal dynamics.

I assumed the highly-upvoted posts would be long-form close readings of the text. In reality, the most liked posts were almost all mashups of "Dream of the Red Chamber + modern scenarios": putting Daiyu's speaking style into workplace communication, or Baochai's social strategies into dormitory relationships. Those who strictly discussed the text were actually in the minority.

I assumed the discussants were similar to me—all students of the same age. But the IP locations showed 60% concentrated in tier-1 and tier-2 cities, and the active users looked more like young professionals and exam candidates. So if the thesis discusses "contemporary reception," the sample profile must be clearly defined; you can't assume it's just the college student demographic.

So I completely changed my topic: from "Causes of the Tragic Love Between Baoyu and Daiyu" to "Modern Dissemination of Character Discourse in Dream of the Red Chamber—Taking Workplace Interpretations of Daiyu and Baochai as Examples." Both are about Dream of the Red Chamber, but the latter is backed by 96 pieces of real data, giving me confidence while writing.

3. AI helps you go from 0 to 1, but you have to walk the path from 1 to 100 yourself

With the data in hand, I let AI continue forward with the second instruction:

Based on these 96 discussion data points about Dream of the Red Chamber, help me analyze: 1. The most discussed themes and high-frequency tags; 2. The most common angles of discussion for each theme; 3. Any findings that contradict intuition; 4. Based on the above analysis, list 5 feasible directions for my academic year paper, noting the data basis for each direction.

One direction AI listed that I hadn't considered: approaching from "The Modern Applicability of Daiyu's Art of Conversation." Its reasoning was that posts tagged with "art of conversation" had overall higher interaction volumes. I never even thought of that angle.

But there are three things it cannot do, and you shouldn't expect it to.

First, judging which direction is worth pursuing. It listed 5 directions, but the judgment that "the data on interpersonal dynamics is sufficient to support a thesis" was something I figured out myself by staring at those 96 entries. It can classify and list, but it cannot judge.

Second, making the final call. Choosing the final direction, structuring the argument, and finding literature were all decided by me, and I still had to pass my advisor's review. AI is an advisor, not a decision-maker. You can mention it in the acknowledgments, but not as a co-author.

Third, you have to go back and read the original work and talk to people. Data tells you what people are discussing, but not why. Why do office workers prefer Baochai? That requires going back to the text or actually interviewing some people. AI can't help with that part of the journey.

4. Final Thoughts

In the past, if a humanities student wanted to choose a topic supported by data, they either had to beg a classmate who could code for help or just give up. Now, all it takes is adding a sentence in Chinese late at night.

I used the saved time to reread the original novel, re-marking the passages where Daiyu retorts sharply. This is how tools should be used: take over the time-consuming tasks, but leave the thinking to yourself.

5. Practical Takeaway: Two Prompts, Copy Directly

Here's the good stuff. Replace the parts in 【】 in the two prompts below with topics relevant to your major, and copy the rest exactly. It works for any major.

Prompt A · Scrape Data

Search Xiaohongshu for posts from the past week, keywords including: 【Keyword 1】【Keyword 2】【Keyword 3】【Keyword 4】. Scrape post content and comments, categorize by 【Dimension 1】【Dimension 2】【Dimension 3】【Dimension 4】. For each entry, extract title, body summary, comment count, like count, publish time, and IP location, add theme tags, and finally tally counts by category.

Prompt B · Analyze

Based on this batch of scraped data, help me analyze: 1. The most discussed themes and high-frequency tags; 2. The most common angles of discussion for each theme; 3. Any findings that contradict intuition; 4. Based on the above analysis, list 5 feasible directions for my 【thesis/assignment/project】, noting the data basis for each direction.

Run through this checklist before starting:

Provide about 4 keywords. Too broad will scrape useless info; too narrow won't yield enough data.

Plan 3-4 classification dimensions in advance. This is the framework you use to find discrepancies; don't categorize randomly.

Limit the time range (e.g., "within one week"). Without limits, old and new posts mix together, making comparison impossible.

Check the count and distribution first, then closely read the high-like and high-comment entries.

User nicknames and avatars appearing in the data must be anonymized before citing them in the thesis.

After running the process, follow this order:

Write down your three "I assumed" statements first, then look for counterexamples in the data. Every counterexample could be a potential topic.

Use Prompt B to get 5 directions, pick only 1-2 with the strongest data backing to proceed with.

Take the direction and chat with teachers and classmates once. Data is one side, human language is the other; both need to align.

Go back to the raw materials (original novels, textbooks, literature) to fill in gaps. AI only handles the 0-to-1 part.

Clearly state the data source, scraping time, and sample size in the thesis; otherwise, it won't hold up.

If you use this method to generate a topic, tell us in the comments how many entries you scraped and which "I assumed" statement was overturned.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts