[quote="gao_yunfan, post:1, topic:1870"]
Just saw this paper on arxiv, the title alone perked me up. Beyond Shapley, an influence-based data auditing pipeline for LLM alignment and evaluation. My advisor recently asked me to read several papers on data filtering, which has been giving me a headache. Everyone knows data quality is important, but quantifying "this conversation is harmful to the model" or "this preference annotation is wrong" has always been vague.
This paper's approach is straightforward: use influence functions instead of Shapley values for data auditing. I had some prior knowledge of Shapley values, a concept from cooperative game theory, calculating...
[/quote]
This paper's approach hits the nail on the head, but influence functions rely on Hessian approximations, which tend to be unstable on large models. I've seen teams force it with EKFAC or low-rank approximations, but computational costs still don't come down. If you're looking at code implementation, check if their appendix mentions tricks like tiny proxy models; actual deployment might depend on them.