Machine Learning

1 readers

1 users here now

Community Rules:

Be nice. No offensive behavior, insults or attacks: we encourage a diverse community in which members feel safe and have a voice.
Make your post clear and comprehensive: posts that lack insight or effort will be removed. (ex: questions which are easily googled)
Beginner or career related questions go elsewhere. This community is focused in discussion of research and new projects that advance the state-of-the-art.
Limit self-promotion. Comments and posts should be first and foremost about topics of interest to ML observers and practitioners. Limited self-promotion is tolerated, but the sub is not here as merely a source for free advertisement. Such posts will be removed at the discretion of the mods.

founded 1 year ago

MODERATORS

communick@academy.garden

[D] Why are ML model outputs not tested regarding statistical significance? (alien.top)

submitted 1 year ago by Tigmib@alien.top to c/machinelearning@academy.garden

57 comments fedilink hide all child comments

Often when I read ML papers the authors compare their results against a benchmark (e.g. using RMSE, accuracy, ...) and say "our results improved with our new method by X%". Nobody makes a significance test if the new method Y outperforms benchmark Z. Is there a reason why? Especially when you break your results down e.g. to the anaylsis of certain classes in object classification this seems important for me. Or do I overlook something?

you are viewing a single comment's thread
view the rest of the comments

[–] Recent_Ad4998@alien.top 1 points 1 year ago (3 children)

One thing I have found is that if you have a large dataset, the standard error can become so small that any difference in average performance will be significant. Obviously not always the case, depending on the size of the variance etc but I imagine it might be why it's often acceptable not to include them.

[–] iswedlvera@alien.top 1 points 1 year ago (1 children)

This is the reason. People do significance tests when you want to draw conclusions with 20 samples on an entire population. If you have thousands of samples there won't be much point.

[–] econ1mods1are1cucks@alien.top 1 points 1 year ago (1 children)

Depends on how big the individual samples are tbh. 1000 samples of 10 people actually sounds like a decent study group

[–] iswedlvera@alien.top 1 points 1 year ago

I see what you mean. Yeah it shouldn't be by default I don't do statistical significance tests.

load more comments (1 replies)