How should I choose a single “best” machine learning model when different models perform best on different OSS projects?

How should I choose a single “best” machine learning model when different models perform best on different OSS projects?

Manage alerts

Loading saved threads...

pritpal singh · External communityPost link
External question — Cross Validated Stack Exchange Author: pritpal singh Original post: https://stats.stackexchange.com/questions/673889 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I am evaluating multiple machine learning classifiers on several open-source software (OSS) projects for solving classification problems. For each project, I train and test all models separately and evaluate their performance using ROC-AUC and PR-AUC. However, I observe that: Different models achieve the best performance on different projects. No single model consistently outperforms others across all projects. Despite this, I need to select one single model to: Perform feature importance analysis, and Serve as a representative model for interpretation and reporting. Given this scenario: What is the principled way to choose a single model? Should I select the model with the best average performance across projects, the most stable model, or the most interpretable model? Are there recommended practices in empirical software engineering or applied ML research for handling this trade-off? On which metric we should rely (ROC-AUC or AUC-PR)? Any guidance or references would be appreciated.
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External question — Cross Validated Stack Exchange Author: pritpal singh Source score (net votes, not local likes): 1 Original post: https://stats.stackexchange.com/questions/673889 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I am evaluating multiple machine learning classifiers on several open-source software (OSS) projects for solving classification problems. For each project, I train and test all models separately and evaluate their performance using ROC-AUC and PR-AUC. However, I observe that: Different models achieve the best performance on different projects. No single model consistently outperforms others across all projects. Despite this, I need to select one single model to: Perform feature importance analysis, and Serve as a representative model for interpretation and reporting. Given this scenario: What is the principled way to choose a single model? Should I select the model with the best average performance across projects, the most stable model, or the most interpretable model? Are there recommended practices in empirical software engineering or applied ML research for handling this trade-off? On which metric we should rely (ROC-AUC or AUC-PR)? Any guidance or references would be appreciated.

Cancel quote

Checking account access…