Abstract:The fair allocation of data revenue is one of the core issues in building sustainable data markets. Compared with traditional production factors, data exhibit several characteristics, such as ex-post value, information asymmetry, costless replication, and externalities, which pose multidimensional challenges for designing revenue allocation strategies. This study focuses on the machine learning model market, which is an important branch of the data market. It systematically reviews the research progress of revenue allocation strategies in this domain, revealing a development trend from homogeneity to differentiation and from short-term to long-term. Specifically, the revenue allocation problem in the machine learning model market is first formalized, and the participants, allocation modes, and objectives are clarified. On this basis, the allocation basis of “homogeneous allocation-differentiated compensation” is organized. In terms of homogeneous contribution measurement, data contribution evaluation methods based on indicators, such as the Shapley value, are summarized. In terms of differentiated compensation, the measurement methods of differentiated indicators such as data cost and data diversity are analyzed, and a hybrid strategy integrating both dimensions is revealed. Furthermore, regarding the dynamic characteristics of the model market over the long term, the impact of strategic behaviors of different participants on revenue allocation and the corresponding response measures are analyzed. Finally, the main challenges in current research are summarized, and future research directions for optimizing revenue allocation strategies are clarified from the perspectives of differentiated compensation and long-term dynamics.