<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>test &#8211; Spencer Greenberg</title>
	<atom:link href="https://www.spencergreenberg.com/tag/test/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.spencergreenberg.com</link>
	<description></description>
	<lastBuildDate>Wed, 17 Dec 2025 01:00:27 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://i0.wp.com/www.spencergreenberg.com/wp-content/uploads/2024/05/cropped-icon.png?fit=32%2C32&#038;ssl=1</url>
	<title>test &#8211; Spencer Greenberg</title>
	<link>https://www.spencergreenberg.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">23753251</site>	<item>
		<title>Is IQ Legitimate or B.S.?</title>
		<link>https://www.spencergreenberg.com/2025/03/is-iq-legitimate-or-b-s/</link>
					<comments>https://www.spencergreenberg.com/2025/03/is-iq-legitimate-or-b-s/#comments</comments>
		
		<dc:creator><![CDATA[Spencer]]></dc:creator>
		<pubDate>Tue, 25 Mar 2025 23:12:21 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[agreeableness]]></category>
		<category><![CDATA[Big Five]]></category>
		<category><![CDATA[cognitive task]]></category>
		<category><![CDATA[extraversion]]></category>
		<category><![CDATA[iq]]></category>
		<category><![CDATA[IQ score]]></category>
		<category><![CDATA[Legitimate]]></category>
		<category><![CDATA[neuroticism]]></category>
		<category><![CDATA[openness]]></category>
		<category><![CDATA[personality]]></category>
		<category><![CDATA[personality test]]></category>
		<category><![CDATA[predicition]]></category>
		<category><![CDATA[predictor]]></category>
		<category><![CDATA[social science]]></category>
		<category><![CDATA[test]]></category>
		<guid isPermaLink="false">https://www.spencergreenberg.com/?p=4365</guid>

					<description><![CDATA[Is the idea of IQ legit or total B.S.? With the replication crisis in social science, it&#8217;s worth asking this since a number of major psychology findings didn&#8217;t hold up under scrutiny. To find out, at Clearer Thinking, we ran a massive study. We tested thousands of people performing random subsets of 62 diverse cognitive [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Is the idea of IQ legit or total B.S.? With the replication crisis in social science, it&#8217;s worth asking this since a number of major psychology findings didn&#8217;t hold up under scrutiny.</p>



<p class="wp-block-paragraph">To find out, at Clearer Thinking, <a href="https://www.clearerthinking.org/post/what-s-really-true-about-intelligence-and-iq-we-empirically-tested-40-claims">we ran a massive study</a>.</p>



<p class="wp-block-paragraph">We tested thousands of people performing random subsets of 62 diverse cognitive tasks (vocab, math, logic, pattern recognition, reaction time, games, memorization, mental rotation, language learning, etc.)</p>



<p class="wp-block-paragraph">We successfully replicated a classic finding: performance on nearly all cognitive tasks correlates positively with performance on the other tasks—a phenomenon known as the &#8220;positive manifold,&#8221; foundational to IQ.</p>



<p class="wp-block-paragraph">IQ scores explained ~45% of variation across our diverse cognitive tasks, aligning with previous research. That&#8217;s very substantial for a single number (IQ score), but also far from capturing everything.</p>



<p class="wp-block-paragraph">Some of the remaining 55% variance is pure noise; the rest likely comes from task skill (developed through practice) and task-specific aptitudes (likely influenced by both genetics and childhood experiences).</p>



<p class="wp-block-paragraph">It&#8217;s also worth noting that while IQ is predictive of a diverse range of intelligence tasks, it doesn&#8217;t necessarily capture ALL that&#8217;s meant by intelligence. It&#8217;s unclear if it includes &#8220;street smarts,&#8221; social skills, or deep nature skills (like hunter-gatherers have).</p>



<p class="wp-block-paragraph">We confirmed IQ predicts interesting outcomes:</p>



<p class="wp-block-paragraph">• Actively open-minded thinking (r=0.43)</p>



<p class="wp-block-paragraph">• Household income (weakly, r=0.15)</p>



<p class="wp-block-paragraph">• Self-reported job performance (for lower IQ ranges only, r=0.48)</p>



<p class="wp-block-paragraph">• Celebrity worship (substantially negatively correlated with IQ, r=-0.42)</p>



<p class="wp-block-paragraph">• No link to happiness or life satisfaction.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph">IQ captures something—but WHAT? There&#8217;s no consensus. Theories include that IQ is&#8230;</p>



<p class="wp-block-paragraph">• A single cognitive resource</p>



<p class="wp-block-paragraph">• Working memory + executive control</p>



<p class="wp-block-paragraph">• A measure of brain integration</p>



<p class="wp-block-paragraph">• The result of overlapping cognitive processes</p>



<p class="wp-block-paragraph">• The top of a hierarchy of abilities</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph">Guess what&#8217;s a BETTER predictor of major life outcomes than IQ?</p>



<p class="wp-block-paragraph">The Big Five personality traits.</p>



<p class="wp-block-paragraph">In our data, openness, conscientiousness, extraversion, agreeableness, and neuroticism (when all 5 are used together) outperformed IQ at predicting most outcomes.</p>



<p class="wp-block-paragraph">IQ clearly matters—it unavoidably jumps out from the data when testing people on a diverse range of cognitive tasks.</p>



<p class="wp-block-paragraph">And yet&#8230;</p>



<p class="wp-block-paragraph">IQ is far from being all that matters and is definitely not destiny.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph">We used the data from our study to create a new cognitive assessment, which analyzes ability across 7 dimensions and aims to provide you with useful insights about your mind.</p>



<p class="wp-block-paragraph">If you&#8217;d like to take it, you can do so here:</p>



<p class="wp-block-paragraph"><a href="https://programs.clearerthinking.org/cognitive-test-intro.html">https://programs.clearerthinking.org/cognitive-test-intro.html</a></p>



<p class="wp-block-paragraph">(Proceeds support our mission.)</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>This piece was first written on March 25, 2025, and first appeared on my website on May 15, 2025.</em></p>



<p class="wp-block-paragraph"></p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2025/03/is-iq-legitimate-or-b-s/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">4365</post-id>	</item>
		<item>
		<title>Testing a Theory Without an Experiment</title>
		<link>https://www.spencergreenberg.com/2018/01/testing-a-theory-without-an-experiment/</link>
					<comments>https://www.spencergreenberg.com/2018/01/testing-a-theory-without-an-experiment/#respond</comments>
		
		<dc:creator><![CDATA[Spencer]]></dc:creator>
		<pubDate>Thu, 25 Jan 2018 13:38:00 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[experiment]]></category>
		<category><![CDATA[odds]]></category>
		<category><![CDATA[predict]]></category>
		<category><![CDATA[test]]></category>
		<guid isPermaLink="false">https://www.spencergreenberg.com/?p=2152</guid>

					<description><![CDATA[You don&#8217;t need to run an experiment to perform a valid test of one of your theories or hypotheses (whether informal or scientific). There is a technique, which I&#8217;ll describe below, that can be far faster, and is used a lot less than it should be (especially when trying to test a theory in science, [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">You don&#8217;t need to run an experiment to perform a valid test of one of your theories or hypotheses (whether informal or scientific). There is a technique, which I&#8217;ll describe below, that can be far faster, and is used a lot less than it should be (especially when trying to test a theory in science, where it could save you an month long experiment, but also, with informal theories in daily life). I aspire to use this approach significantly more often than I do now.</p>



<p class="wp-block-paragraph"><strong>How to Test a Theory Without Doing An Experiment</strong>:</p>



<p class="wp-block-paragraph">Think of as many things as you can that are predicted by your theory, and try to identify at least one prediction that has the following four properties:</p>



<p class="wp-block-paragraph">(1) your theory implies it is very likely to be true.</p>



<p class="wp-block-paragraph">(2) it seems very unlikely to be true if your theory is false (i.e., there is no reason you can see to make that prediction other than if you believe in your theory). You may want to ask others if they agree that the prediction seems very unlikely to be true from the perspective of those who don&#8217;t believe your theory; otherwise, you might be deceiving yourself.</p>



<p class="wp-block-paragraph">(3) you don&#8217;t already know the answer to whether it is actually true or not (you are merely guessing it is true because it is implied by your theory). This is important because if you already know the answer, you may be tricking yourself one way or another (e.g., by having already used this evidence in the construction of your theory, or by choosing this particular test because on some level you know it will make your theory look good).</p>



<p class="wp-block-paragraph">(4) you can reliably look up whether that prediction is true or not. Then all you have to do is look up whether the prediction is true!</p>



<p class="wp-block-paragraph">If it turns out to be true, that provides evidence to support your theory. More specifically, the more likely that true result is given your theory, and the less likely that true result is (to someone who doesn&#8217;t know your theory or if your theory is invalid), the more evidence the test provides. To explore the math for only a moment: that test of your theory should change your prior odds (i.e., your estimated probability that your theory is true, divided by the probability that it is not true) by using what&#8217;s known as the &#8220;Bayes factor,&#8221; namely, your estimation probability. Looking up the answer to that question would yield the result it did, if your theory is true divided by the probability it would yield that result if your theory is not true. </p>



<p class="wp-block-paragraph">In other words:</p>



<p class="wp-block-paragraph">odds your theory is true = ( probability of that result if your theory is true ÷ probability of that result if your theory is false) × prior odds your theory is true from before the test</p>



<p class="wp-block-paragraph">Odds can be confusing to think about (remember that it is the probability that something is true divided by the probability that it is false). For instance, odds of 2 would mean that a thing is twice as likely to be the case then it is to not be the case.</p>



<p class="wp-block-paragraph"></p>



<p class="wp-block-paragraph">It&#8217;s worth noting that the above procedure can work really effectively to test your hypotheses for yourself, but it&#8217;s NOT a good way to provide evidence that your theory is true to other people that don&#8217;t trust you (i.e., your honest pursuit of truth <strong>or</strong> your competence). The reason is that it&#8217;s easy to claim you conducted this fair test of your theory, when in fact, you already knew the answer to how the test would turn out (and therefore were incentivized to choose this particular test of your theory rather than another). Or perhaps you interpreted your theory in a weird way or tweaked it after the fact to make it look like it correctly predicted the outcome when it actually didn&#8217;t. Or maybe you screwed up one of the steps accidentally. In other words, the evidence this procedure produces is only trustworthy if you really trust the rigor and character of the person who produced the evidence.</p>



<p class="wp-block-paragraph">This is why experiments are, at least in theory (but unfortunately too often not in practice), a better way to gather evidence when you need to communicate those results to others who don&#8217;t trust you, especially if you pre-register your study and analysis plan, and if you release your materials and the data, you collect.</p>



<p class="wp-block-paragraph">If you&#8217;re interested in the concepts underpinning the approach I&#8217;ve described here, you may want to check out our program on Bayesian thinking (which you may or may not be happy to know uses almost no math):</p>



<p class="wp-block-paragraph"><a href="http://programs.clearerthinking.org/question_of_evidence.html" target="_blank" rel="noreferrer noopener">http://programs.clearerthinking.org/question_of_evidence.ht…</a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2018/01/testing-a-theory-without-an-experiment/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">2152</post-id>	</item>
		<item>
		<title>Are you a &#8220;credentialist&#8221; or &#8220;non-credentialist&#8221;?</title>
		<link>https://www.spencergreenberg.com/2017/07/are-you-a-credentialist-or-non-credentialist/</link>
					<comments>https://www.spencergreenberg.com/2017/07/are-you-a-credentialist-or-non-credentialist/#comments</comments>
		
		<dc:creator><![CDATA[Spencer]]></dc:creator>
		<pubDate>Fri, 21 Jul 2017 23:01:00 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[attitude]]></category>
		<category><![CDATA[credentialist]]></category>
		<category><![CDATA[credentials]]></category>
		<category><![CDATA[degrees]]></category>
		<category><![CDATA[expert]]></category>
		<category><![CDATA[influence]]></category>
		<category><![CDATA[non-credentialist]]></category>
		<category><![CDATA[opinions]]></category>
		<category><![CDATA[professional]]></category>
		<category><![CDATA[questions]]></category>
		<category><![CDATA[scale]]></category>
		<category><![CDATA[spectrum]]></category>
		<category><![CDATA[subject]]></category>
		<category><![CDATA[test]]></category>
		<category><![CDATA[trait]]></category>
		<category><![CDATA[trust]]></category>
		<guid isPermaLink="false">https://www.spencergreenberg.com/?p=4394</guid>

					<description><![CDATA[Are you a &#8220;credentialist&#8221; or &#8220;non-credentialist&#8221;? Here&#8217;s a test I designed so that you can find out. After noticing a number of times that people&#8217;s feelings about formal credentials can differ dramatically and that this seems to impact their views on certain important topics, I&#8217;ve been working on defining a &#8220;credentialist&#8221; trait (or attitude). In [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Are you a &#8220;credentialist&#8221; or &#8220;non-credentialist&#8221;? Here&#8217;s a test I designed so that you can find out.</p>



<p class="wp-block-paragraph">After noticing a number of times that people&#8217;s feelings about formal credentials can differ dramatically and that this seems to impact their views on certain important topics, I&#8217;ve been working on defining a &#8220;credentialist&#8221; trait (or attitude).</p>



<p class="wp-block-paragraph">In a nutshell, the non-credentialist/credentialist spectrum, as I&#8217;m defining it, captures how important a person thinks formal credentials are, as well as how they feel those credentials should influence who we should trust and who should express opinions (e.g., should only formalized experts comment on a topic, or is it good for non-experts to comment as well?)</p>



<p class="wp-block-paragraph">I developed a 4-minute test to measure the trait, so if you&#8217;d like to find out if you are a &#8220;credentialist&#8221; or &#8220;non-credentialist&#8221; (which I define as being in the top or bottom 20th percentile of each trait) or find out your own credentialist score, you can take the test here:</p>



<p class="wp-block-paragraph"><a href="https://programs.clearerthinking.org/credentialist_test.html">https://programs.clearerthinking.org/credentialist_test.html</a></p>



<p class="wp-block-paragraph">Here are simplified/extreme prototypes to illustrate the distinction (few people are as extreme as these prototypes):</p>



<p class="wp-block-paragraph">*Credentialists* get annoyed when someone without the right credential is giving their opinion on a topic, are impressed by formal degrees (e.g. PhDs and MDs), do not like it when non-experts have their own personal theory about a topic, trust people a lot more when they have formal credentials, think its unlikely someone could get really good at a complex topic without formal training, think that non-experts should not contradict experts, would go to school to learn a topic if they wanted to get good at it, find it annoying if startup founders talk about disrupting industries they have not already worked in, tend to describe people in terms of their schooling and job history (rather than, for example, their personality), and trust the opinion&#8217;s of people a lot more if they went to an excellent college.</p>



<p class="wp-block-paragraph">*Non-credentialists* think that it fine (or even good) to express opinions when you&#8217;re a non-expert, are not particularly impressed by formal degrees, don&#8217;t view degrees or certifications as a strong indicator of trust, think it&#8217;s fine (or even good) for non-experts to criticize the views of experts, are prone to teach themselves material rather than going to school for it, don&#8217;t mind startup founders attempting to disrupt industries from the outside, tend not to describe people in terms of their schooling and job history, and don&#8217;t view the quality of the college a person went to as a significant factor in whether to trust their opinions.</p>



<p class="wp-block-paragraph">I measured the trait on a 0 to 1 scale, and on the 143 people I collected it for, it has a mean of 0.55 and a standard deviation of 0.16, with a pretty nice bell curve shape.</p>



<p class="wp-block-paragraph">Interestingly, I found a very low correlation between credentialist scores and education [r=0.03], identifying as female [r=-0.04], and income [r=0.03], and little correlation with age [r=-.08]. This suggests that credentialist differences have little to do with demographic characteristics!</p>



<p class="wp-block-paragraph">Furthermore, whether you are a credentialist does not even seem to have that much to do with whether you yourself have credentials, as responses to the question &#8220;I myself have substantial formal credentials&#8221; had a correlation of only about r=0.13 with the credentialist scores. It also has little to do with how ambitious people are, as responses to &#8220;I have highly ambitious goals for what I will achieve in my life&#8221; had a correlation of only r=0.07 to credentialist scores.</p>



<p class="wp-block-paragraph">Of the 22 questions I tested people&#8217;s agreement on in order to measure the trait, the 2 most effective questions (in the sense that they correlate highly with the average of the other questions but don&#8217;t correlate very highly with each other) are:</p>



<p class="wp-block-paragraph">Q1: It annoys me when someone without the right credential is giving their opinion on a topic (e.g., a non-doctor commenting about medicine or a non-accountant commenting on accounting)</p>



<p class="wp-block-paragraph">[r=0.71 against the average of the other 21 questions]</p>



<p class="wp-block-paragraph">and</p>



<p class="wp-block-paragraph">Q2: I am very impressed by formal degrees (e.g., PhDs, MDs, JDs, etc.)</p>



<p class="wp-block-paragraph">[r=0.54 against the average of the other 21 questions, yet a fairly low r=0.24 with respect to Q1]</p>



<p class="wp-block-paragraph">The 22 questions I developed hang together nicely and point in generally the same direction. Basic factor analysis revealed only one main factor in the questions, and the question least correlated to the others still had a positive correlation of r=0.34 with the average of the other questions (which was &#8220;If I heard that someone had won a prize in their field, I would think very highly of it&#8221;).</p>



<p class="wp-block-paragraph">Most people fall in the middle of this trait, of course (e.g., viewing credentials as at least somewhat positive but the lack of them not that negative), without an extreme viewpoint either way. However, here are some anonymized qualitative responses I collected from people at the tail ends of the spectrum:</p>



<p class="wp-block-paragraph">Credentialists:</p>



<p class="wp-block-paragraph">&#8220;I don&#8217;t respect people who don&#8217;t have formal credentials, and one of my biggest pet peeves is people&#8211;even really intelligent people&#8211;speaking about things they aren&#8217;t experts on. Just because someone is known, or even renowned in another field, doesn&#8217;t place their opinion on another topic anywhere above another non-educated person.&#8221;</p>



<p class="wp-block-paragraph">&#8220;People with formal credentials have generally gone through a rigorous peer review process that demands a considerable depth of understanding and knowledge.&#8221;</p>



<p class="wp-block-paragraph">&#8220;People who don&#8217;t have credentials have no business talking about things they know nothing about. That&#8217;s how misinformation gets spread, and misinformation is harmful to society as a whole.&#8221;</p>



<p class="wp-block-paragraph">Non-credentialists:</p>



<p class="wp-block-paragraph">&#8220;I think that people can have a valid voice no matter their level of formal schooling. The opposite also holds true: a degree is not necessarily representative of ability.&#8221;</p>



<p class="wp-block-paragraph">&#8220;I feel that colleges have become scams. They are money sponges that make you pay for your own brainwashing. I value intelligence and how well-read someone is on a subject, and we don&#8217;t need a self-appointed team of left-wing experts to &#8216;allow&#8217; us to do that for ourselves anymore. People with credentials have more money than sense and have been taught what NOT to think more than they have been taught HOW to think.&#8221;</p>



<p class="wp-block-paragraph">&#8220;While I respect those who have worked to earn professional credentials, people who are self-taught oftentimes know much more about a subject than someone with an expensive degree.&#8221;</p>



<p class="wp-block-paragraph">Note: designing this scale and writing this post is a very non-credentialist thing to do since I&#8217;m not a social scientist.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>This piece was first written on July 21, 2017, and first appeared on my website on June 10, 2025.</em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2017/07/are-you-a-credentialist-or-non-credentialist/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">4394</post-id>	</item>
		<item>
		<title>Testing Too Many Hypotheses</title>
		<link>https://www.spencergreenberg.com/2011/10/testing-too-many-hypotheses/</link>
					<comments>https://www.spencergreenberg.com/2011/10/testing-too-many-hypotheses/#comments</comments>
		
		<dc:creator><![CDATA[Spencer]]></dc:creator>
		<pubDate>Mon, 10 Oct 2011 17:16:40 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[experiments]]></category>
		<category><![CDATA[hypotheses]]></category>
		<category><![CDATA[hypothesis test]]></category>
		<category><![CDATA[p-values]]></category>
		<category><![CDATA[probability]]></category>
		<category><![CDATA[science]]></category>
		<category><![CDATA[statistics]]></category>
		<category><![CDATA[test]]></category>
		<guid isPermaLink="false">http://www.spencergreenberg.com/?p=249</guid>

					<description><![CDATA[For each dataset, there is a limit to what we can use that dataset to test. Using the standard p-value based methods of science, the more hypotheses we check against the data, the more likely it will be that some of these checks give inaccurate conclusions. And this presents a big problem for the way [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>For each dataset, there is a limit to what we can use that dataset to test. Using the standard <a href="http://en.wikipedia.org/wiki/P-value">p-value based methods</a> of science, the more hypotheses we check against the data, the more likely it will be that some of these checks give inaccurate conclusions. And this presents a big problem for the way science is practiced.</p>
<p>Let&#8217;s take an example to illustrate the principle. Suppose that you have information about 1000 people selected at random from the U.S. adult population. Your dataset includes these people&#8217;s heights, weights, ages, shoe sizes, and so forth. Now, if your goal is to know the mean height of all people in America, you can produce an estimate of this quantity by averaging the heights of the 1000 people you have information about. Despite the fact that your sample contains just 1000 people, rather than the full set of 230,000,000 or so American adults of interest, your estimate will, with high probability, be within a couple of inches of the total population mean height. This is due to the fact that the 1000 people were sampled at random (so we shouldn&#8217;t expect our sample to differ from the entire population in a systematic way) and because the standard deviation of heights is not very large (if there were tremendous outliers in the data, such as 500 foot tall giants, we would need more samples to get an accurate estimate). This idea is made precise by the central limit theorem. It tells us how likely the true entire population mean is to fall different distances from our sample estimate, and says that the error of our estimate decreases like one over the square root of the size of our sample.</p>
<p>The same technique could work to approximate the mean weight of adult Americans, or age, or shoe size, or number of children. And in each case, the estimate would, with high probability, be quite accurate. We could even, if we liked, estimate all these quantities simultaneously if we collected all of this information about each of our 1000 people. But the more quantities we estimate, the greater the chance that at least one estimate is quite inaccurate. Since each estimate has some chance of being bad, if we make a sufficiently large number of estimates we should expect to get unlucky at some point and end up with one or more bad ones. So, if we aren&#8217;t just estimating mean height, but rather the mean of 50 different traits, we cannot claim that all 50 of these estimate are likely to be good. We should expect that some of them will be inaccurate, though we don&#8217;t know which ones.</p>
<p>This is where problems arise. Suppose that you are a researcher who is trying to find interesting differences between, say, southerners and northerners in the United States. Your dataset of 1000 adults contains 500 people from each group. What do you do? Well, it might seem reasonable to go ahead and compute the mean value of many different traits, and look at how these means differ between the two groups, to see if you can find any large differences that seem interesting. For instance, you may compute the average salary of each group, and see if they deviate from each other by a large enough amount to be deemed statistically significant. If they don&#8217;t, you can try another trait like IQ, or number of children, and repeat the process. If you try enough different traits, hopefully you&#8217;ll eventually find an intriguingly large difference between the groups.</p>
<p>The trouble is, we know that if you estimate a large number of quantities, some of them will be inaccurate, and so some of the apparent differences between your two groups may just be due to these inaccuracies. If you test enough traits, you will eventually find differences between the populations that look significant, even though it is just the result of chance.</p>
<p>In fact, even if northerners and southerners had no systematic differences between them, there would still be apparent differences that arose just from the particular sample of 1000 people you happened to have data on. For example, in your dataset, it just might happen that the northerners have lower numbers of children than southerners, even if this isn&#8217;t true for the underlying populations of all northerners and southerners. If you were to publish this finding, without making mention of the number of hypotheses you tested before finding it, it may seem that you had produced a meaningful result. In fact, the assessment of this result should take into account the number of hypotheses (e.g. northerners have smaller shoe sizes than southerners, northerners have greater salaries than southerners, etc.) that you tested before you discovered this one (and the <a href="http://en.wikipedia.org/wiki/Multiple_comparisons">p-values can be modified to include this information</a>). The most significant seeming deviation between the groups found after testing 100 different hypotheses is very likely greatly inflated by chance. Whereas if you had only tested a small number of hypotheses against your data, and found a strong result, this would likely be a meaningful finding.</p>
<p>As a general rule, the greater the number of data points you have, the larger the number of quantities you can accurately estimate from your dataset. On a set of just 10 points, you may not even be able to get an accurate estimate of the mean value of a single trait (unless the trait had very slow standard deviation). Whereas on a dataset of a billion points, you probably could estimate dozens of quantities accurately.</p>
<p>Unfortunately, when you&#8217;re reading a paper, there is no way to tell how many hypotheses the researcher tested on his dataset unless he chooses to publish it. And there is a strong incentive to obscure this information. If a researcher releases the fact that he tested 20 hypotheses before finding 1 which was statistically significant, readers may discredit the result, or reviewers may reject it for publication. And if the researcher spent a lot of time and money collecting his dataset, it would feel like a waste to give up on the data just because his first five hypotheses tested on it don&#8217;t pan out. It might take a lot of restraint to not just keep testing hypothesis after hypothesis until he finds something publishable.</p>
<p>But even if researchers were excessively careful, that wouldn&#8217;t fully resolve the problem. When a hypothesis is confirmed by a dataset, we must consider whether it is truly a confirmation of the hypothesis being tested, or a result of the fact that 20 researchers tested 20 false hypotheses, and this one of the 20 happened to seem true by chance. That is, if enough hypotheses are tested over all, we may find a large number of false hypotheses among them that just happen to seem true.</p>
<p>What makes this problem more pernicious is that when a hypothesis fails to pan out, the result is often not published. This is due to the fact that hypothesis disconfirmations (e.g. &#8220;no association was found between cabbage eating and longevity&#8221;) are generally less interesting and harder to publish than confirmations (e.g. &#8220;an association was found between cabbage eating and longevity&#8221;). But since most new hypotheses in science turn out to be false, we should expect the number of negative results to be very large (except in situations where previously well validated results are being confirmed). Hence, the number of published test results will be much less than the number of total tests conducted, with test failures substantially underreported. So there is no good way to tell how many times a hypotheses failed to be confirmed by tests before one researcher finally ran one that seemed to confirm it. And if a very large number of false hypotheses are tested, but mostly just the ones that turn out to look true are published, you could end up with a field&#8217;s journals being flooded with false but seemingly verified hypotheses. In exploratory fields where almost all hypotheses are false, and where disconfirmations of a hypothesis are almost never published, you might even get into a situation where <a href="http://www.plosmedicine.org/article/info:doi/10.1371/journal.pmed.0020124">most published research findings are false</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2011/10/testing-too-many-hypotheses/feed/</wfw:commentRss>
			<slash:comments>6</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">249</post-id>	</item>
	</channel>
</rss>
