Data Audits: Finally, Scrutiny Directed at Training Data Itself
Community Discussion · Policy

Data Audits: Finally, Scrutiny Directed at Training Data Itself

Can't Finish Reading PapersCan't Finish Reading PapersJul 282026/07/28 64 views

Just came across this on arXiv, and the title alone perked me right up. "Beyond Shapley," an impact-based data audit pipeline for LLM alignment and evaluation. My advisor recently asked me to read a few papers on data filtering, and my head is spinning. Everyone knows data quality matters, but quantifying things like "this conversation is harmful to the model" or "this preference label is wrong" has always been vague.

1 replies

?
Ctrl + Enter to reply
Tiangong
TiangongJul 30(edited)

[quote="gao_yunfan, post:1, topic:1870"]

Just saw this paper on arxiv, the title alone perked me up. Beyond Shapley, an influence-based data auditing pipeline for LLM alignment and evaluation. My advisor recently asked me to read several papers on data filtering, which has been giving me a headache. Everyone knows data quality is important, but quantifying "this conversation is harmful to the model" or "this preference annotation is wrong" has always been vague.

This paper's approach is straightforward: use influence functions instead of Shapley values for data auditing. I had some prior knowledge of Shapley values, a concept from cooperative game theory, calculating...

[/quote]

This paper's approach hits the nail on the head, but influence functions rely on Hessian approximations, which tend to be unstable on large models. I've seen teams force it with EKFAC or low-rank approximations, but computational costs still don't come down. If you're looking at code implementation, check if their appendix mentions tricks like tiny proxy models; actual deployment might depend on them.