Community Discussion · Tracks

Weekend Deep Dive into Chery Sales Data with Python Scrapers: Hit Many Pitfalls

SlippageSlippageAug 22026/08/02 197 views

Over the weekend, I wanted to pull sales data for Chery's various brands to check the delivery curve for the Luxeed V9. I figured since this is public data, writing a Python script would take a few minutes. Ended up taking nearly two hours.

Let me explain what I was trying to do and why someone involved in quantitative trading cares about this. At most, I think the authenticity and timeliness of sales data are more reliable than any analyst report. I watch order book data on my trading terminal every day, so I'm used to grabbing raw data directly and don't trust second-hand info.

But if beginners jump straight in, they'll likely get stuck at the first step—finding the data structure.

Step 1: Figure Out Where the Data Is

Open the IT Home news page for Chery's July sales. The table on the page is clear:

  • Chery: 183,615 units in July (domestic), +28% YoY
  • Exports: 202,533 units, +70.1% YoY
  • New Energy Vehicles (NEVs): 129,067 units, +97.5% YoY

But if you right-click "View Page Source," you'll find the table data is plain text, not structured JSON. You have to parse it manually.

Biggest Pitfall for Beginners: Copy-pasting directly into Excel and finding the format messed up. Because numbers contain commas (183,615), Excel treats them as two separate numbers. I tried it; wasted time.

Step 2: Write a Simple Scraper

I'm used to using Python's requests library to fetch pages and BeautifulSoup to parse them. But honestly, for beginners, the most reliable method is manual copying.

I wrote a snippet of code, roughly:

Then I discovered a problem: IT Home's page structure changes frequently. Last week's structure differed from this week's. Your hardcoded find conditions might fail next week.

Real-Life Pitfall Story: I initially tried regex matching, but it matched the "Related News" list at the bottom of the page instead of the sales data. Wasted 15 minutes.

Step 3: Manually Verify the Data

Quants have a habit: Never trust a single data source. I pulled data from several sources simultaneously:

  • Gasgoo: Fengyun A9 pre-orders hit 31,000 in 15 days, with users under 35 accounting for 70%
  • Netcar: Exports broke 200,000 for the first time
  • Sohu: ROE 36.5%, ranked 30th globally

Conclusion: Chery's July domestic sales were 183k units, exports 202k, totaling 385k, with NEVs at 129k. Three sources aligned in methodology, making the data credible.

Step 4: Convert Data into Analyzable Format

For beginners, I recommend using an Excel spreadsheet formatted like this:

Brand July Sales YoY Change Key Model
Chery (Domestic) 183,615 +28% Fengyun A9
Exports 202,533 +70.1% —
NEVs 129,067 +97.5% All
Luxeed 10,000+ +227.7% V9

Note: Luxeed V9 delivered over 10,000 units in a single month, a 227.7% YoY increase. This number seems exaggerated, but the source material explicitly states it. I checked, and the Luxeed V9 is a high-end NEV MPV with a high unit price, so the base sales volume is small, making growth rates easy to inflate.

An Important Quant Mindset: High growth rate does not equal large absolute value; look at the base. A 227.7% YoY increase—if last year's sales were only 3,000 and this year's are 10,000, that's impressive. But if last year's were 5,000 and this year's are 16,000, the growth rate drops. I suggest beginners calculate compound growth rates for MoM and YoY rather than just looking at single-month figures.

Step 5: Trend Assessment

After running the data, I reached a conclusion: Chery's export business is becoming a new growth engine. Single-month exports of 200k, up 70% YoY, mark the first time a Chinese automaker has broken this threshold.

On the NEV front, breaking 100k units for four consecutive months, along with 31k pre-orders for the Fengyun A9 in 15 days, indicates real demand in the 150k RMB pure-electric sedan market. However, caution is needed regarding the Luxeed V9's 227.7% growth—the high-end MPV market capacity is limited, and sustainability depends on next month's data.

What to try after learning this: Compare sales curves over the past 12 months, draw a K-line chart, and observe seasonal fluctuations. Or pull peer data (BYD, Geely, Changan) for horizontal comparison. You could even write a simple script to automatically pull data weekly and update your database.

Anyway, having data in hand keeps the mind at ease.

1 replies

?
Ctrl + Enter to reply
Compliance Anxiety

This crawler approach is pretty good, but when building risk control models, inconsistent data sources are the biggest nightmare. Regarding the 227% growth rate you mentioned, did you cross-validate it with month-over-month changes or inventory turnover rates to prevent single-month data from being manipulated?