Using data in discussions: datasets are useful as reference points when someone claims something unusual. "I have not seen that reported in the data" is different from "that is impossible", but data gives you something to say.
Coming back to: A community side-effect dataset, with its response rate and biases
post #13 is right about the mechanism and I think understates the practical bit.
Bias toward positive outcomes: datasets collected by members are biased toward people who found the compounds useful. People who did not respond do not return. People who had bad outcomes might have left the community.
I disagree with the reply above, and I think the disagreement is substantive rather than terminological.
The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.
post #31 is right about the mechanism and I think understates the practical bit.
Temporal bias: older data in a dataset might reflect conditions (supplier, formulation, context) that have changed. Newer data is more current.
Picking up post #44: that is the part I would want checked first.
Collection methods: ask how the data were collected. Longitudinal tracking over months is stronger than retrospective recall. Prospective measurement (done while experiencing something) is stronger than memory afterward.
Limitations of datasets: all community-collected data has limitations. The population is self-selected (people in this community are not representative of all people using these compounds). Reporting bias is real (remarkable outcomes get reported; mundane outcomes do not).
Read the full topic (102 posts)
This topic was referenced in
- Contributing data without breaching anyone's privacyData & Tools › Datasets · 71 replies
Suggested topics
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
About the Datasets category
Community-collected datasets, their collection methods, and their limitations. This post is a community wiki: any member at trust level 3 or above can edit it, and every edit is recorded with its author and a…
|
+6 | 11 | 42k | 12mo |
|
Whether an aggregate is worth publishing at all: a disputed topic — does this still hold?
Whether an aggregate is worth publishing at all: a disputed topic — does this still hold? I have a specific reason for asking rather than idle curiosity, and the context is below. I would like to understand…
|
+14 | 18 | 3.4k | 2mo |
|
Cleaning a self-reported dataset and what you throw away — what changed since
Posting this under the heading it deserves: Cleaning a self-reported dataset and what you throw away — what changed since Everything below is what sits behind that. Posting the method first, because I know…
|
2 | 2k | 1h | |
|
Whether an aggregate is worth publishing at all: a disputed topic
On the subject in the title: Whether an aggregate is worth publishing at all: a disputed topic Working notes rather than a conclusion. I would like to understand what this number means before I repeat it…
|
+20 | 24 | 18k | 13mo |
|
Contributing data without breaching anyone's privacy — what changed since
On the subject in the title: Contributing data without breaching anyone's privacy — what changed since Working notes rather than a conclusion. A documentation question rather than an analytical one. I have a…
|
2 | 58k | 17mo |
Related topics — sharing the tags observational data, site feedback, data table
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Impurity thresholds: where the common numbers come from
Impurity thresholds: where the common numbers come from — setting out what I have, and where I think it stops being reliable. A documentation question rather than an analytical one. I have a certificate in…
|
+11 | 15 | 825 | 2d |
|
Area percent versus weight percent: the confusion that causes most arguments
Posting this under the heading it deserves: Area percent versus weight percent: the confusion that causes most arguments Everything below is what sits behind that. Posting the method first, because I know…
|
+10 | 14 | 2.1k | 2mo |
|
Tag proliferation and whether we should prune
On the subject in the title: Tag proliferation and whether we should prune Working notes rather than a conclusion. Process question about how this site works. I have read the guidelines and the trust-level…
|
2 | 11k | 11mo | |
|
[2026 update] Why 99.2% and 97.8% on the same lot can both be correct
Why 99.2% and 97.8% on the same lot can both be correct — that is the question, and I have not found it answered plainly anywhere I have looked. I would like to understand what this number means before I…
|
+1 | 5 | 719 | 2d |
|
A friction point in the promotion workflow
A friction point in the promotion workflow — setting out what I have, and where I think it stops being reliable. Nominating something for promotion into the documentation commons. The topic in question keeps…
|
+4 | 8 | 491 | 11mo |