An analysis of data sets used to train and validate cost prediction systems

Title	An analysis of data sets used to train and validate cost prediction systems
Author(s)	Mair, M., Shepperd, M. and Jørgensen, M.
Details	Conference Proceedings: 2005
Abstract	OBJECTIVE – to build up a picture of the nature and type of data sets being used to develop and evaluate different software project effort prediction systems. We believe this to be important since there is a growing body of published work that seeks to assess different prediction approaches. METHOD – we performed an exhaustive search from 1980 onwards from three software engineering journals for research papers that used project data sets to compare cost prediction systems. RESULTS – this identified a total of 50 papers that used, one or more times, a total of 71 unique project data sets. We observed that some of the better known and easily accessible data sets were used repeatedly making them potentially disproportionately influential. Such data sets also tend to be amongst the oldest with potential problems of obsolescence. We also note that only about 60% of all data sets are in the public domain. Finally, extracting relevant information from research papers has been time consuming due to different styles of presentation and levels of contextural information. CONCLUSIONS – first, the community needs to consider the quality and appropriateness of the data set being utilised; not all data sets are equal. Second, we need to assess the way results are presented in order to facilitate meta-analysis and whether a standard protocol would be appropriate.
DOI	http://doi.acm.org/10.1145/1082983.1083166
BibTex	View Citation
Topics	Application, Cost Estimation, Data Sets, Mapping Study, Secondary Study