<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>probability &#8211; Spencer Greenberg</title>
	<atom:link href="https://www.spencergreenberg.com/tag/probability/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.spencergreenberg.com</link>
	<description></description>
	<lastBuildDate>Sat, 30 May 2026 23:12:36 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.0</generator>

<image>
	<url>https://i0.wp.com/www.spencergreenberg.com/wp-content/uploads/2024/05/cropped-icon.png?fit=32%2C32&#038;ssl=1</url>
	<title>probability &#8211; Spencer Greenberg</title>
	<link>https://www.spencergreenberg.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">23753251</site>	<item>
		<title>How to spot real expertise</title>
		<link>https://www.spencergreenberg.com/2024/04/how-to-spot-real-expertise/</link>
					<comments>https://www.spencergreenberg.com/2024/04/how-to-spot-real-expertise/#respond</comments>
		
		<dc:creator><![CDATA[admin]]></dc:creator>
		<pubDate>Tue, 23 Apr 2024 13:33:35 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[belief]]></category>
		<category><![CDATA[calibration]]></category>
		<category><![CDATA[consensus]]></category>
		<category><![CDATA[epistemic humility]]></category>
		<category><![CDATA[epistemics]]></category>
		<category><![CDATA[evaluating evidence]]></category>
		<category><![CDATA[evidence]]></category>
		<category><![CDATA[expertise]]></category>
		<category><![CDATA[honesty]]></category>
		<category><![CDATA[humility]]></category>
		<category><![CDATA[intellectual humility]]></category>
		<category><![CDATA[judgment]]></category>
		<category><![CDATA[knowledge]]></category>
		<category><![CDATA[probability]]></category>
		<category><![CDATA[rationality]]></category>
		<category><![CDATA[scout mindset]]></category>
		<category><![CDATA[steelman]]></category>
		<category><![CDATA[steelmanning]]></category>
		<category><![CDATA[strawman]]></category>
		<category><![CDATA[uncertainty]]></category>
		<category><![CDATA[updating]]></category>
		<guid isPermaLink="false">https://www.spencergreenberg.com/?p=3902</guid>

					<description><![CDATA[Thanks go to Travis (from the Clearer Thinking team) for coauthoring this with me. This is a cross-post from Clearer Thinking. How can you tell who is a valid expert, and who is full of B.S.? On almost any topic of importance you can find a mix of valid experts (who are giving you reliable [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph"><em>Thanks go to Travis (from the Clearer Thinking team) for coauthoring this with me.</em> <em>This is a cross-post from <a href="https://www.clearerthinking.org/post/how-to-spot-real-expertise?utm_source=ClearerThinking.org&amp;utm_campaign=a6a0ff049e-EMAIL_CAMPAIGN_FAKE_EXPERTISE&amp;utm_medium=email&amp;utm_term=0_f2e9d15594-b71c1a1f3d-%5BLIST_EMAIL_ID%5D&amp;mc_cid=a6a0ff049e&amp;mc_eid=dea552ccde">Clearer Thinking</a>. </em></p>



<p class="wp-block-paragraph" id="viewer-6ho89124">How can you tell who is a valid expert, and who is full of B.S.?</p>



<p class="wp-block-paragraph" id="viewer-toa9l129">On almost any topic of importance you can find a mix of valid experts (who are giving you reliable information) and false but confident-seeming &#8220;experts&#8221; (who are giving you misinformation). To make matters even more confusing, sometimes the fake experts even have very impressive credentials, and every once in a while, the real, genuine experts are entirely self-taught.</p>



<p class="wp-block-paragraph" id="viewer-nh6hz132">Here are 12 signs we look for in an expert to help us determine whether they are trustworthy.&nbsp;</p>



<h2 class="wp-block-heading" id="viewer-c5pf3134">1. They have deep factual knowledge</h2>



<p class="wp-block-paragraph" id="viewer-u4tmf136">Let’s start with the obvious: for most topics, a lot of factual knowledge is required before you can have genuine expertise. This means that a genuine expert will have an impressive command of the relevant (non-debated) facts on the topic of their expertise. Thankfully, it&#8217;s a lot easier to tell if an expert has a strong command of the non-debated facts than whether they are correct about more controversial claims.&nbsp;</p>



<h2 class="wp-block-heading" id="viewer-9rkmj138">2. They communicate their confidence levels</h2>



<p class="wp-block-paragraph" id="viewer-ikf4r140">Not all knowledge is equally well-established. Even theories that are widely accepted enjoy different levels of support from the relevant evidence. When an expert regularly pretends that all their claims are equally well-established, they demonstrate they are willing to make you believe something is certain when it isn&#8217;t.</p>



<p class="wp-block-paragraph" id="viewer-99oed142">It’s a good sign that someone treats their subject with the nuance expected from genuine expertise, when they indicate how confident they are (e.g., “It&#8217;s been shown in many high-quality studies that…”, or “My best guess is…”), and they explain limitations in the evidence they are using (e.g., “this is unfortunately based on just one study, but that is all that currently exists”)</p>



<h2 class="wp-block-heading" id="viewer-hoas5144">3. They admit not knowing</h2>



<p class="wp-block-paragraph" id="viewer-z5138146">Genuine experts also sometimes say that they don’t know the answer to a question, or that the answer is generally not known by anyone. This is important because every topic will have some unknowns, and no expert can know everything about a topic. Telling you when they don&#8217;t know is a sign that, when they say they <em>do</em>&nbsp;know, they actually do know.</p>



<h2 class="wp-block-heading" id="viewer-3bw8y150">4. They tell you to look at sources other than themselves</h2>



<p class="wp-block-paragraph" id="viewer-868ro152">This might happen when an expert doesn’t know the answer to a question, or when they want to help you go beyond the answer they can give you. Genuine experts don&#8217;t seek to be seen as a sole arbiter of knowledge or authority on a topic (which can be an indication that ego, rather than truth-seeking, is a primary motivation for them), but instead encourage you to look at resources other than the ones they have produced.</p>



<h2 class="wp-block-heading" id="viewer-9grt2154">5. They use logic and evidence</h2>



<p class="wp-block-paragraph" id="viewer-wqk8m156">Anyone can use rhetorical devices like emotional appeals, no matter how wrong they are, but a well-reasoned argument that uses valid logic and strong evidence will tend to point toward truth. Or, put another way, using strong logic and strong evidence is easier to do when you&#8217;re right, whereas emotional appeals are no easier when you&#8217;re right than when you&#8217;re wrong.&nbsp;&nbsp;</p>



<h2 class="wp-block-heading" id="viewer-89qm4158">6. They cite high-quality evidence</h2>



<p class="wp-block-paragraph" id="viewer-98ifh160">Some evidence is much more reliable than other evidence, and those who rely on the less reliable kinds when the more reliable kinds exist probably aren&#8217;t doing the best job they can at figuring out the truth. For this reason, genuine experts cite high-quality evidence when it exists (e.g., looking at multiple randomized controlled trials for causal claims) rather than low-quality evidence (e.g., just talking about personal anecdotes), and when high-quality evidence doesn’t exist, they cite the highest quality evidence that does exist.</p>



<h2 class="wp-block-heading" id="viewer-xi60e162">7. They acknowledge the consensus</h2>



<p class="wp-block-paragraph" id="viewer-auj7t164">Consensus views among experts are more often correct than the idiosyncratic views of just one or two experts. The consensus will not always be right, of course, but often it will be the best understanding we have available. That’s why reliable experts are transparent about the degree to which their opinion differs from the majority of experts, provide reasoned explanations for any deviations, and they are cautious not to present fringe theories as mainstream. This shows a deep engagement with the topic of their expertise and also an adherence to ethical standards of honesty and accuracy in communication.</p>



<h2 class="wp-block-heading" id="viewer-oky7e166">8. They change their mind</h2>



<p class="wp-block-paragraph" id="viewer-6w5gt168">Genuine experts will change their minds about topics within their expertise in response to evidence and arguments. It’s hard to become an expert in something without having been wrong from time-to-time.</p>



<p class="wp-block-paragraph" id="viewer-xy6s3170">That means that anyone claiming to be an expert who has never changed their mind probably has not found and corrected their mistakes. Relatedly, changing one&#8217;s mind in response to evidence is also a sign of the epistemic humility associated with genuine expertise.</p>



<p class="wp-block-paragraph" id="viewer-h4l3f172">Of course, if someone has a long history of being wrong, that is evidence against them being a genuine expert, not in favor of it. But, since everyone makes some mistakes, if they make mistakes from time to time and then note they were wrong and improve their beliefs, that is a sign that they are following the evidence where it leads rather than continuing to believe what they do regardless of the evidence.</p>



<h2 class="wp-block-heading" id="viewer-94wkg174">9. They Steelman</h2>



<p class="wp-block-paragraph" id="viewer-1edoz176">When you ‘straw man’ an argument, you misrepresent or oversimplify someone else&#8217;s position to make it easier to attack or refute. Instead of dealing with the actual argument, you replace it with a weaker version that distorts the original point, which you then argue against. The opposite of this is called ‘steelmanning’, and it involves presenting the strongest possible version of an argument you’re objecting to, even if it&#8217;s more robust than the one originally presented. This approach aims to strengthen the opposing case in order to facilitate a more genuine and constructive debate.&nbsp;</p>



<p class="wp-block-paragraph" id="viewer-kib65178">The most reliable experts will accurately present the strongest arguments made by those that disagree with them while pointing out flaws in those arguments, rather than focusing on just weak arguments from the other side or just mocking the other side (including ad hominem attacks rather than focusing on the substance of the claims of the other side). This is important because knocking down a weak argument from the other side of a debate does little to show the other side is wrong; you have to refute the strongest claims of the other side to actually show they are wrong. Additionally, demonstrating a knowledge of the strongest arguments against your own position shows a deeper level of expertise than only understanding the opposing point of view at a superficial level.</p>



<h2 class="wp-block-heading" id="viewer-mmo40180">10. They clearly explain their reasons for believing</h2>



<p class="wp-block-paragraph" id="viewer-yh6ch182">The philosopher Daniel Dennett <a target="_blank" href="https://books.google.co.uk/books?id=C5pUnN1-vhcC&amp;pg=PT16&amp;dq=%22if+I+can%E2%80%99t+explain+something+I%E2%80%99m+doing+to+a+group+of+bright+undergraduates,+I+don%E2%80%99t+really+understand+it+myself%22&amp;redir_esc=y#v=onepage&amp;q=%22if%20I%20can%E2%80%99t%20explain%20something%20I%E2%80%99m%20doing%20to%20a%20group%20of%20bright%20undergraduates%2C%20I%20don%E2%80%99t%20really%20understand%20it%20myself%22&amp;f=false" rel="noreferrer noopener">has said</a>: “if I can’t explain something I’m doing to a group of bright undergraduates, I don’t really understand it myself.” This sentiment is echoed by philosopher John Searle, who said “In general, I feel if you can&#8217;t say it clearly you don&#8217;t understand it yourself.”&nbsp;</p>



<p class="wp-block-paragraph" id="viewer-spi6a186">When communicating with non-experts, genuine experts are often able to give clear, easy-to-follow (and, ideally, checkable) explanations for why they believe what they believe &#8211; without dumbing down the points. They avoid unnecessary jargon and technical language (which sounds smart but makes their arguments very difficult for their audience to follow). Not every genuine expert is able to do this, but the ability to do this well is a sign of genuine expertise. This is important because an expert who cannot explain their ideas clearly will end up requiring you to believe them based on their authority rather than engaging with the arguments themselves. And sometimes, people claiming to be experts will hide behind technical expertise and jargon so that you won&#8217;t notice that their arguments are actually weak.</p>



<h2 class="wp-block-heading" id="viewer-gs759188">11. They have a track record</h2>



<p class="wp-block-paragraph" id="viewer-11d0m190">Sometimes genuine experts will have track records of predictions or successes that you can check, and this provides direct evidence of their knowledge or skill. Unfortunately, this only applies to some fields, like chess masters, martial experts who fight in tournaments, experts who make public predictions about the economy or politics, etc.</p>



<h2 class="wp-block-heading" id="viewer-pzysj192">12. They use multiple lenses</h2>



<p class="wp-block-paragraph" id="viewer-o5ipy194">The world is complex and multi-faceted, and any one simple theory is going to fail to explain a lot of what&#8217;s really going on. For this reason, genuine experts tend to look at problems from multiple frames and perspectives; they don&#8217;t act as though one way of looking at things solves all problems, or that one solution works for all problems, or that one simple theory explains everything.</p>



<p class="wp-block-paragraph" id="viewer-b0ax0196">So the next time you hear claims from an alleged expert on a topic that is important to you, you may want to consider: how many of these signs of expertise do they exhibit? You can use this checklist, considering if they:</p>



<ol class="wp-block-list">
<li>have deep factual knowledge</li>



<li>communicate their confidence levels</li>



<li>admit not knowing</li>



<li>tell you to look at sources other than themselves</li>



<li>use logic and evidence</li>



<li>cite high-quality evidence</li>



<li>acknowledge the consensus</li>



<li>change their mind</li>



<li>steelman</li>



<li>clearly explain their reasons for believing</li>



<li>have a track record</li>



<li>use multiple lenses</li>
</ol>



<p class="wp-block-paragraph" id="viewer-6liwv235">And if you’re seeking to be an expert in something yourself, you may want to ask yourself: “to what extent do I exhibit these traits?”Being able to discern genuine expertise from B.S. requires good judgment. If you’d like to improve your skills at making accurate judgments, why not try our <a href="https://www.openphilanthropy.org/calibration">Calibrate Your Judgment tool,</a> created in partnership with Open Philanthropy.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>This piece first appeared on Clearer Thinking.org on April 16, 2024, and first appeared on my website on April 22, 2024.</em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2024/04/how-to-spot-real-expertise/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">3902</post-id>	</item>
		<item>
		<title>Did That Treatment Actually Help You?</title>
		<link>https://www.spencergreenberg.com/2023/04/did-that-treatment-actually-help-you/</link>
					<comments>https://www.spencergreenberg.com/2023/04/did-that-treatment-actually-help-you/#respond</comments>
		
		<dc:creator><![CDATA[admin]]></dc:creator>
		<pubDate>Sat, 15 Apr 2023 22:18:00 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[Bayes]]></category>
		<category><![CDATA[Bayesian]]></category>
		<category><![CDATA[beliefs]]></category>
		<category><![CDATA[causality]]></category>
		<category><![CDATA[estimaton]]></category>
		<category><![CDATA[experiments]]></category>
		<category><![CDATA[placebo]]></category>
		<category><![CDATA[probability]]></category>
		<category><![CDATA[scientific method]]></category>
		<category><![CDATA[updating]]></category>
		<guid isPermaLink="false">https://www.spencergreenberg.com/?p=3539</guid>

					<description><![CDATA[A mistake we all make sometimes is attributing an improvement to whatever we&#8217;ve tried recently. For instance, we may get medicine from a doctor (or go to an acupuncturist) and feel better, so we conclude it worked. But did it actually work, or was it just chance? Here&#8217;s a trick to help you decide: What [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">A mistake we all make sometimes is attributing an improvement to whatever we&#8217;ve tried recently. For instance, we may get medicine from a doctor (or go to an acupuncturist) and feel better, so we conclude it worked. But did it actually work, or was it just chance? Here&#8217;s a trick to help you decide:</p>



<p class="wp-block-paragraph">What matters (evidence-wise) is how likely that level of improvement would have been in that time period if the treatment works relative to how likely that improvement would have been if the treatment is useless.</p>



<p class="wp-block-paragraph">For something like tiredness, which tends to fluctuate a lot, feeling somewhat less tired than normal after two weeks may provide almost no evidence a treatment worked. But if you feel less tired than you have in 10 years, that could be strong evidence!</p>



<p class="wp-block-paragraph">To give another example, if you&#8217;ve had a rash without a break for years, and the rash goes away in one day with a new cream, that is very strong evidence the cream worked. But if the rash very often comes and goes on its own, or it took six months of using the cream before it disappeared, its disappearance provides little evidence of effectiveness.</p>



<p class="wp-block-paragraph">More formally, the amount of evidence an improvement gives you (in favor of the treatment working) is:</p>



<p class="wp-block-paragraph">Bayes Factor = the probability that you&#8217;d see this level of improvement given that the treatment works / the probability that you&#8217;d see this level of improvement given that the treatment doesn&#8217;t work</p>



<p class="wp-block-paragraph">In words, this is just &#8220;how many times more likely is it that you&#8217;d see this level of improvement during this period of time if the treatment works compared to if it doesn&#8217;t work.&#8221;</p>



<p class="wp-block-paragraph">This Bayes Factor is what you multiply your prior odds by. So if, before trying the treatment, you thought there were 1 to 3 odds of it working (i.e., a 25% chance), and if you now you get a Bayes factor of 6, you should now believe there are 6*(1/3) = 2 to 1 odds that it works (i.e., a 66% chance).</p>



<p class="wp-block-paragraph">While it&#8217;s rare to be able to do this calculation precisely, it&#8217;s this general way of thinking (in terms of relative likelihoods, comparing a world where the treatment works to one where it doesn&#8217;t) that&#8217;s important. I find this to be an especially helpful application of Bayes&#8217; rule which can guide practical decision-making (e.g., whether to stick with a new treatment).</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>This piece was first written on April 15, 2023, and first appeared on this site on August 2, 2023.</em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2023/04/did-that-treatment-actually-help-you/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">3539</post-id>	</item>
		<item>
		<title>Demystifying p-values</title>
		<link>https://www.spencergreenberg.com/2022/12/demystifying-p-values/</link>
					<comments>https://www.spencergreenberg.com/2022/12/demystifying-p-values/#comments</comments>
		
		<dc:creator><![CDATA[admin]]></dc:creator>
		<pubDate>Sat, 31 Dec 2022 20:40:00 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[alpha]]></category>
		<category><![CDATA[alternative hypothesis]]></category>
		<category><![CDATA[Bayesianism]]></category>
		<category><![CDATA[false positives]]></category>
		<category><![CDATA[frequentism]]></category>
		<category><![CDATA[garden of forking paths]]></category>
		<category><![CDATA[multiple hypothesis testing]]></category>
		<category><![CDATA[null hypothesis]]></category>
		<category><![CDATA[null hypothesis significance testing]]></category>
		<category><![CDATA[p-hacking]]></category>
		<category><![CDATA[p-values]]></category>
		<category><![CDATA[probability]]></category>
		<category><![CDATA[publication bias]]></category>
		<category><![CDATA[random chance]]></category>
		<category><![CDATA[replication crisis]]></category>
		<category><![CDATA[statistical significance]]></category>
		<category><![CDATA[statistics]]></category>
		<category><![CDATA[underpowered]]></category>
		<guid isPermaLink="false">https://www.spencergreenberg.com/?p=3382</guid>

					<description><![CDATA[There is a tremendous amount of confusion around what a p-value actually is, despite their widespread use in science. Here is my attempt to explain the concept of p-values concisely and clearly (including why they are useful and what often goes wrong with them). — What&#8217;s a p-value? — If you run a study, then [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">There is a tremendous amount of confusion around what a p-value actually is, despite their widespread use in science. Here is my attempt to explain the concept of p-values concisely and clearly (including why they are useful and what often goes wrong with them).</p>



<p class="wp-block-paragraph"><strong>— What&#8217;s a p-value? —</strong></p>



<p class="wp-block-paragraph">If you run a study, then (all else equal, aside from rare edge cases) the lower the p-value, the lower the chance that your results are due to random chance or luck.</p>



<p class="wp-block-paragraph">More precisely: a p-value is the probability you&#8217;d get a result at least as extreme as what you got IF there were actually no effect (or if some other pre-specified &#8220;null hypothesis&#8221; is true).</p>



<p class="wp-block-paragraph">So it&#8217;s a probability calculated based on assuming that there is no effect (or assuming that a pre-specified &#8220;null hypothesis&#8221; is true). Here the phrase &#8220;no effect&#8221; would mean, in the case of a study on a new medicine, that the medicine doesn&#8217;t do anything.</p>



<p class="wp-block-paragraph">To put it in terms of coin flips: suppose you&#8217;re trying to decide if a coin is fair (i.e., if it has an equal chance of landing on heads and tails &#8211; so that&#8217;s your &#8220;null hypothesis&#8221; in this context). You flip the coin 100 times and get 60 heads. You calculate the p-value (p=0.06).</p>



<p class="wp-block-paragraph">This p-value tells you there&#8217;s a 6% chance you&#8217;d get 60 or more heads OR 60 or more tails out of 100 flips if the coin were actually fair.</p>



<p class="wp-block-paragraph">What makes p-values useful is that when they are high, you usually can&#8217;t rule out your effect being due to random chance or luck. And, when they are very low, random chance is (in most cases) unlikely to be the explanation for your result.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><strong>— What&#8217;s the problem with p-values? —</strong></p>



<p class="wp-block-paragraph">In social science, p&lt;0.05 is often used as the cutoff for a &#8220;successful&#8221; result (i.e., they treat the effect as real and potentially publishable). This is an arbitrary cutoff; there&#8217;s nothing special about 0.05. The phrase &#8220;statistically significant&#8221; is defined simply to mean that p&lt;0.05.</p>



<p class="wp-block-paragraph">There are many ways that p-values get commonly misused, creating lots of problems. For instance:</p>



<p class="wp-block-paragraph">• p-values often get misinterpreted as the probability that an effect is not real (recall: p-values are actually the probability of getting a result this extreme if there is no effect, which is not the same thing)</p>



<p class="wp-block-paragraph">• If you see one study where the main finding&#8217;s p-value is, say, 0.05, and another study where the main finding&#8217;s p-value is, say, 0.01, it&#8217;s tempting to conclude that the finding of the 2nd study is much less likely to be the result of chance (e.g., 1/5th as likely) than the 1st study&#8217;s finding. Unfortunately, we can&#8217;t draw this conclusion. The probability that a study&#8217;s finding is the result of chance is not the same as the p-value, and in fact, it can&#8217;t even be calculated just by knowing the p-value.</p>



<p class="wp-block-paragraph">• Because a p-value threshold is often used for a result to be publishable (p&lt;0.05 in social science), researchers sometimes engage in fishy methods to get their p-values below the threshold. This is known as &#8220;p-hacking.:</p>



<p class="wp-block-paragraph">• A result&#8217;s p-value (or &#8220;statistical significance&#8221;) is sometimes focused on instead of focusing on other factors that are also important. For instance, a result may have a low p-value but be such a weak effect that it&#8217;s totally useless or uninteresting.</p>



<p class="wp-block-paragraph">• While a low p-value helps you rule out the possibility that your effect is merely due to random chance, unfortunately, that&#8217;s all it helps you with. But researchers sometimes act as though it tells them more than that. Even an extremely low p-value doesn&#8217;t mean an effect is &#8220;real&#8221; or that the effect means what you think. Low p-values can result from a variety of causes, including mistakes in experimental design or confounds.</p>



<p class="wp-block-paragraph">Here&#8217;s another way to think about what a p-value is and isn&#8217;t that some people find helpful: a p-value does not tell you the probability that your result is due to chance. It tells you how consistent your results are with being due to chance. (I&#8217;m paraphrasing from <a href="https://statmodeling.stat.columbia.edu/2013/03/12/misunderstanding-the-p-value/#comment-143473">here</a>.) So, the lower the p-value, the less consistent your results are with them being due to chance.</p>



<p class="wp-block-paragraph">It&#8217;s interesting to note that, empirically, results with lower p-values are more likely to be genuine effects (i.e., not false positives). I looked at results for 325 psychology study replications, and when the original study p-value was at most 0.01, about 72% replicated. When p&gt;0.01, only 48% did.</p>



<p class="wp-block-paragraph">Ultimately, p-values are a useful (though often abused) statistical tool.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><strong>— BONUS APPENDIX: what&#8217;s the chance of a hypothesis being &#8220;true&#8221; if p&lt;0.05?  —</strong></p>



<p class="wp-block-paragraph">One annoying thing about p-values is that they don&#8217;t answer the question we are usually interested in. Usually, we want to know something like &#8220;What&#8217;s the probability that my hypothesis is true?&#8221; or &#8220;What&#8217;s the probability that the effect of this drug is bigger than X?&#8221; but p-values don&#8217;t tell us those things.</p>



<p class="wp-block-paragraph">However, we can put a different spin on p-values to get them to answer questions that are closer to what we&#8217;re really interested in. Let&#8217;s think of p-values as giving us a decision procedure (in an overly simplified world where you either &#8220;believe&#8221; in an effect or you fail to believe in it).&nbsp;</p>



<p class="wp-block-paragraph">Suppose you test 100 totally separate, previously unexplored hypotheses about humans, and suppose that you commit to &#8220;believe&#8221; a hypothesis is true if and only if you get p&lt;0.05 (and otherwise, you don&#8217;t believe it).</p>



<p class="wp-block-paragraph">I think it&#8217;s realistic that in a social science context, most hypotheses studied will be false since discovering novel, publishable hypotheses about humans is hard. So let&#8217;s suppose that 80% of the hypotheses you test are *not* true.&nbsp;</p>



<p class="wp-block-paragraph">Finally, suppose that you use a large enough number of participants in your studies so that if you are testing for the presence of a real effect, there is an 80% chance you&#8217;ll be able to find it (this 80% figure is a common recommendation for &#8220;statistical power&#8221;).&nbsp;</p>



<p class="wp-block-paragraph">Under these assumptions, if you test 100 hypotheses, then you will end up believing in 20 hypotheses, and 80% of those you believe will be true (with the other 20% being false positives). That means that of the results you believe in, 80% will be correct! Of course, this assumes no mistakes are made in the process of designing the experiment, running the statistics, and so on.</p>



<p class="wp-block-paragraph">Here&#8217;s how the math works out if you&#8217;re curious:</p>



<p class="wp-block-paragraph">• Out of the 100 hypotheses, 20 will be true, and of those, you&#8217;ll believe 16 = 0.80 * 20 (these are the true positives) and fail to believe 4 (these are the false negatives).</p>



<p class="wp-block-paragraph">• Out of the 100 hypotheses, 80 will be false, and of those, you&#8217;ll believe 4 = 0.05 * 80 (these are the false positives), and you&#8217;ll reject 76 (these are the true negatives).</p>



<p class="wp-block-paragraph">Of course, if the numbers here had been different, the conclusions would be different as well. For instance, imagine if you started with 2000 hypotheses, and this time, imagine that only 1% of them were true. If the power was still 80%, then:</p>



<p class="wp-block-paragraph">&nbsp;• Out of the 2000 hypotheses, 20 of them would be true, and of those, you&#8217;d believe 16 (0.80 * 20) of them (these are true positives) and fail to believe 4 of them (these are false negatives).</p>



<p class="wp-block-paragraph">• Out of the 2000 hypotheses, 1980 would be false, and of those, you&#8217;d believe 99 (0.05*1980) of them (these are false positives), and you&#8217;d reject the other 1881 of them (these are true negatives).</p>



<p class="wp-block-paragraph">• So, altogether, you&#8217;d believe 115 (16 + 99) hypotheses, of which only 16 would&#8217;ve actually been true, so of the results you believe in, less than 14% would be correct!&nbsp;</p>



<p class="wp-block-paragraph">From analyses like these, we can see that the probability that a specific hypothesis is true, given that we&#8217;ve found p&lt;0.05, depends on a variety of factors, including the sample size, the true effect size, the base rate probability that a new hypothesis tested by that researcher is true, the probability of errors being made in the experimental design or statistical analysis, and so on.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph">In real life:</p>



<p class="wp-block-paragraph">(1) Studies often don&#8217;t use large enough numbers of participants (and so are underpowered).</p>



<p class="wp-block-paragraph">(2) Researchers sometimes engage in p-hacking to artificially lower their p-values to help their papers get published.</p>



<p class="wp-block-paragraph">(3) Researchers often don&#8217;t carefully track how many hypotheses they&#8217;ve really tested.</p>



<p class="wp-block-paragraph">(4) The decision procedure described above is often not adhered to so strictly (e.g., a result of p=0.08 might be treated as suggestive evidence for the hypothesis, and hence the hypothesis is not rejected).</p>



<p class="wp-block-paragraph">(5) Real hypotheses often have auxiliary assumptions beyond what the p-value accounts for (such as an assumption that there is a lack of confounders, a lack of serious errors in the experimental setup, and so on).</p>



<p class="wp-block-paragraph">I personally don&#8217;t like thinking in terms of this decision procedure for p-values because I believe that modeling hypotheses as &#8220;true&#8221; or &#8220;false&#8221; is not a good approach to thinking clearly. This is because I believe it&#8217;s usually much better to think in terms of probabilities rather than a &#8220;true&#8221;/&#8221;false&#8221; dichotomy when trying to understand the answers to complex questions.</p>



<p class="wp-block-paragraph">Some people have argued that we should switch to a Bayesian approach to hypothesis testing since such an approach avoids many of the issues of p-values (including avoiding the problematic &#8220;true&#8221;/&#8221;false&#8221; dichotomy). But it also introduces other challenges, such as how to come up with an appropriate &#8220;prior&#8221; (which represents one&#8217;s belief about the probability of the hypothesis having different strengths of effects prior to seeing the study results).</p>



<p class="wp-block-paragraph"></p>



<p class="wp-block-paragraph"><em>This piece was first written on December 31, 2022, and first appeared on this site on April 2, 2023.</em></p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><a href="https://www.guidedtrack.com/programs/4zle8q9/run?essaySpecifier=%3A+Demystifying%20p-values" target="_blank" rel="noreferrer noopener">If you read this line, please do us a favor and click here to answer one quick question.</a></p>



<p class="wp-block-paragraph"></p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2022/12/demystifying-p-values/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">3382</post-id>	</item>
		<item>
		<title>Importance Hacking: a major (yet rarely-discussed) problem in science</title>
		<link>https://www.spencergreenberg.com/2022/12/importance-hacking-a-major-yet-rarely-discussed-problem-in-science/</link>
					<comments>https://www.spencergreenberg.com/2022/12/importance-hacking-a-major-yet-rarely-discussed-problem-in-science/#comments</comments>
		
		<dc:creator><![CDATA[admin]]></dc:creator>
		<pubDate>Tue, 20 Dec 2022 01:45:00 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[beauty hacking]]></category>
		<category><![CDATA[career incentives]]></category>
		<category><![CDATA[chance]]></category>
		<category><![CDATA[clarity]]></category>
		<category><![CDATA[culture of science]]></category>
		<category><![CDATA[fraud]]></category>
		<category><![CDATA[generalizability]]></category>
		<category><![CDATA[generalizability crisis]]></category>
		<category><![CDATA[hacking]]></category>
		<category><![CDATA[honesty]]></category>
		<category><![CDATA[importance hacking]]></category>
		<category><![CDATA[incentives]]></category>
		<category><![CDATA[integrity]]></category>
		<category><![CDATA[novelty hacking]]></category>
		<category><![CDATA[open science]]></category>
		<category><![CDATA[overclaiming]]></category>
		<category><![CDATA[p-hacking]]></category>
		<category><![CDATA[probability]]></category>
		<category><![CDATA[psychological science]]></category>
		<category><![CDATA[publish or perish]]></category>
		<category><![CDATA[reasoning processes]]></category>
		<category><![CDATA[replication crisis]]></category>
		<category><![CDATA[science]]></category>
		<category><![CDATA[social science]]></category>
		<category><![CDATA[statistics]]></category>
		<category><![CDATA[usefulness hacking]]></category>
		<category><![CDATA[veracity]]></category>
		<guid isPermaLink="false">https://www.spencergreenberg.com/?p=3057</guid>

					<description><![CDATA[I first published this post on the Clearer Thinking blog on December 19, 2022, and first cross-posted it to this site on January 21, 2023. You have probably heard the phrase &#8220;replication crisis.&#8221; It refers to the grim fact that, in a number of fields of science, when researchers attempt to replicate previously published studies, [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph"><em>I first published this post on the <a href="https://www.clearerthinking.org/post/importance-hacking-a-major-yet-rarely-discussed-problem-in-science">Clearer Thinking blog</a> on December 19, 2022, and first cross-posted it to this site on January 21, 2023.</em></p>



<p class="wp-block-paragraph" id="viewer-1d12a"></p>



<p class="wp-block-paragraph" id="viewer-104ln">You have probably heard the phrase &#8220;replication crisis.&#8221; It refers to the grim fact that, in a number of fields of science, when researchers attempt to replicate previously published studies, they fairly often don&#8217;t get the same results. The magnitude of the problem depends on the field, but in psychology, it seems that something like <a rel="noreferrer noopener" href="http://datacolada.org/47" target="_blank"><u>40% of studies in top journals</u></a> don&#8217;t replicate. We&#8217;ve been tackling this crisis with our new <a rel="noreferrer noopener" href="https://replications.clearerthinking.org/" target="_blank"><u><em>Transparent Replications</em></u></a> project, and this post explains one of our key ideas.</p>



<p class="wp-block-paragraph" id="viewer-2dn5g">Replication failures are sometimes simply due to bad luck, but more often, they are caused by p-hacking &#8211; the use of fishy statistical techniques that lead to statistically significant (but misleading or erroneous) results. As big a problem as p-hacking is, there is another substantial problem in science that gets talked about much less. Although certain subtypes of this problem have been named previously, to my knowledge, the problem itself has no name, so I&#8217;m giving it one: &#8220;Importance Hacking.&#8221;</p>



<p class="wp-block-paragraph" id="viewer-3hoev">Academics want to publish in the top journals in their field. To understand Importance Hacking, let&#8217;s consider a (slightly oversimplified) list of the three most commonly-discussed ways to get a paper published in top psychology journals:</p>



<ol class="wp-block-list">
<li><strong>Conduct valuable research</strong> &#8211; make a genuinely interesting or important discovery, or add something valuable to the state of scientific knowledge. This is, of course, what just about everyone wants to do, but it&#8217;s very, very hard!</li>



<li><strong>Commit fraud</strong> &#8211; for instance, by making up your data. Thankfully, very few people are willing to do this because it&#8217;s so unethical. So this is by far the least used approach.</li>



<li><strong>p-hack</strong> &#8211; use fishy statistics, HARKing (i.e., hypothesizing after the results are known), selective reporting, using hidden <a href="https://en.wikipedia.org/wiki/Researcher_degrees_of_freedom" target="_blank" rel="noreferrer noopener"><u>researcher degrees of freedom</u></a>, etc., in order to get a p&lt;0.05 result that is actually just a false positive. This is a major problem and the focus of the replication crisis. Of course, false positives can also come about without fault, due to bad luck.</li>
</ol>



<p class="wp-block-paragraph" id="viewer-5plkf">But here is a fourth way to get a paper published in a top journal: Importance Hacking.</p>



<p class="wp-block-paragraph" id="viewer-ctrs5">4. <strong>Importance Hack</strong> &#8211; get a result that is actually not interesting, not important, and not valuable, but write about it in such a way that reviewers are convinced it is interesting, important, and/or valuable, so that it gets published.</p>



<p class="wp-block-paragraph" id="viewer-f54g1">For research to be valuable to society (and, in an ideal world, publishable in top journals), it must be true AND interesting (or important, useful, etc.). Researchers sometimes p-hack their results to skirt around the &#8220;true&#8221; criterion (by generating interesting false positives). On the other hand, Importance Hacking is a method for skirting the &#8220;interesting&#8221; criterion.</p>



<p class="wp-block-paragraph" id="viewer-ft7mi">Importance Hacking is related to concepts like <em>hype</em> and <em>overselling</em>, though hype and overselling are far more general. Importance Hacking refers specifically to a phenomenon whereby research with little to no value gets published in top journals due to the use of strategies that lead reviewers to misinterpret the work. On the other hand, hype and overselling are used in many ways in many stages of research (including to make valuable research appear even more valuable).</p>



<p class="wp-block-paragraph" id="viewer-dd0l9">One way to understand importance hacking is by comparing it to p-hacking. P-hacking refers to a set of bad research practices that enable researchers to publish non-existent effects. In other words, p-hacking misleads paper reviewers into thinking that non-existent effects are real. Importance Hacking, on the other hand, encompasses a different set of bad research practices: those that lead paper reviewers to believe that real (i.e., existent) results that have little to no value actually have substantial value.</p>



<p class="wp-block-paragraph" id="viewer-2tioa">This diagram illustrates how I think Importance Hacking interferes with the pipeline of producing valuable research:</p>



<figure class="wp-block-image"><img data-recalc-dims="1" decoding="async" src="https://i0.wp.com/static.wixstatic.com/media/f4e552_e1a60b1c65514edf9fef562a77c5c4ba~mv2.jpg/v1/fill/w_1480%2Ch_904%2Cal_c%2Cq_85%2Cusm_0.66_1.00_0.01%2Cenc_auto/f4e552_e1a60b1c65514edf9fef562a77c5c4ba~mv2.jpg?w=750&#038;ssl=1" alt=""/></figure>



<p class="wp-block-paragraph" id="viewer-7u47q">There are a number of subtypes of Importance Hacking based on the method used to make a result appear interesting/important/valuable when it&#8217;s not. Here is how I subdivide them:</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading" id="viewer-brv18"></h2>



<h2 class="wp-block-heading" id="viewer-fh6np">Types of Importance Hacking</h2>



<p class="wp-block-paragraph" id="viewer-a5mla"><strong>1. Hacking Conclusions:</strong> make it seem like you showed some interesting thing X but actually show something else (X′) which sounds similar to X but is much less interesting/important. In these cases, researchers do not truly find what they imply they have found. This phenomenon is also closely connected with validity issues.</p>



<ul class="wp-block-list">
<li><em>Example 1: showing X is true in a simple video game but claiming that X is true in real life.</em></li>



<li><em>Example 2: showing A and B are correlated and claiming that A causes B (when really A and B are probably both caused by some third factor C, which makes the finding much less interesting).</em></li>



<li><em>Example 3: if a researcher claims to be measuring “aggression,” and couches all conclusions in these terms but is actually measuring milliliters of hot sauce that a person puts in someone else&#8217;s food. Their result about aggression will be valid only insofar as it is true that this is a valid measure of aggression.</em></li>



<li>Example 4: some types of hacking conclusions would fall under the terms &#8220;overclaiming&#8221; or &#8220;overgeneralizing;&#8221; Tal Yarkoni has a relevant paper called <a href="https://mzettersten.github.io/assets/pdf/ManyBabies_BBS_commentary.pdf" target="_blank" rel="noreferrer noopener"><em><u>The Generalizability Crisis</u></em></a><em>.</em></li>
</ul>



<p class="wp-block-paragraph" id="viewer-365fm"><strong>2. Hacking Novelty: </strong>refer to something in a way that makes it seem more novel or unintuitive than it is. Perhaps the result is already well known or is merely what just about everyone&#8217;s common sense would already tell them is true. In these cases, researchers really do find what they claim to have found, but what they found is not novel (despite them making it seem so). Hacking Novelty is also connected to the &#8220;Jingle-jangle&#8221; fallacy &#8211; where people can be led to believe two identical concepts are different because they have different names (or, more subtly, because they are operationalized somewhat differently).</p>



<ul class="wp-block-list">
<li><em>Example 1: showing something that is already well-known but giving it a new name that leads people to think it is something new. The concept of “grit” has received this criticism; some people claim it could turn out to be just another word for conscientiousness (or already known facets of conscientiousness) &#8211; though this question does not yet seem to be settled (different sides of this debate can be found in these papers: </em><a rel="noreferrer noopener" href="https://www.researchgate.net/publication/6290064_Grit_Perseverance_and_Passion_for_Long-Term_Goals" target="_blank"><em><u>1</u></em></a><em>, </em><a rel="noreferrer noopener" href="https://journals.sagepub.com/doi/pdf/10.1002/per.2171" target="_blank"><em><u>2</u></em></a><em>, </em><a rel="noreferrer noopener" href="https://drive.google.com/file/d/1NzMPCgZ_Ipbmzewgaj0dmopkfLq582NA/view" target="_blank"><em><u>3</u></em></a><em> and <u><a href="https://www.researchgate.net/publication/304032119_Much_Ado_About_Grit_A_Meta-Analytic_Synthesis_of_the_Grit_Literature">4</a></u>).</em></li>



<li><em>Example 2: showing that A and B are correlated, which seems surprising given how the constructs are named, but if you were to dig into how A and B were measured, it would be obvious they would be correlated.</em></li>



<li><em>Example 3: showing a common-sense result that almost everyone already would predict but making it seem like it&#8217;s not obvious (e.g., by giving it a fancy scientific name).</em></li>
</ul>



<p class="wp-block-paragraph" id="viewer-a209k"><strong>3. Hacking Usefulness: </strong>make a result seem useful or relevant to some important outcome when in fact, it&#8217;s useless and irrelevant. In these cases, researchers find what they claim to have found, but what they find is not useful (despite them making it sound useful).</p>



<ul class="wp-block-list">
<li><em>Example: focusing on statistical significance when the effect size is so small that the result is useless. Clinicians often distinguish between “statistical significance” and “clinical significance” to highlight the pitfalls of ignoring effect sizes when considering the importance of a finding.</em></li>
</ul>



<p class="wp-block-paragraph" id="viewer-etfss"><strong>4. Hacking Beauty: </strong>make a result seem clean and beautiful when in fact, it&#8217;s messy or hard to interpret. In these cases, researchers focus on certain details or results and tell a story around those, but they could have focused on other details or results that would have made the story less pretty, less clear-cut, or harder to make sense of. This is related to Giner-Sorolla’s 2012 paper <a href="https://journals.sagepub.com/doi/pdf/10.1177/1745691612457576" target="_blank" rel="noreferrer noopener"><em><u>Science or art: How aesthetic standards grease the way through the publication bottleneck but undermine science</u></em></a><em>. </em>Hacking beauty sometimes reduces to selective reporting of some kind (i.e., selective reporting of measures, analyses, or studies) or at least of selective focus on certain findings and not others. This becomes more difficult with pre-registration; if you have to report the results of planned analyses, there’s less room to make them look pretty (you could just <em>say</em> they’re pretty, but that seems like overclaiming)</p>



<ul class="wp-block-list">
<li><em>Example: emphasizing the parts of the result that tell a clean story while not including (or burying somewhere in the paper) the parts that contradict that story</em></li>
</ul>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph" id="viewer-56mr8">Science faces multiple challenges. Over the past decade, the <a rel="noreferrer noopener" href="https://en.wikipedia.org/wiki/Replication_crisis" target="_blank"><u>replication crisis</u></a> and subsequent <a rel="noreferrer noopener" href="https://en.wikipedia.org/wiki/Open_science" target="_blank"><u>open science movement</u></a> have greatly increased awareness of p-hacking as a problem. Measures have begun to be put in place to reduce p-hacking. Importance Hacking is another substantial problem, but it has received far less attention.</p>



<figure class="wp-block-image"><img data-recalc-dims="1" decoding="async" src="https://i0.wp.com/static.wixstatic.com/media/f4e552_94289803042f43d68a85e7c490b1fa1c~mv2.jpg/v1/fill/w_1480%2Ch_1110%2Cal_c%2Cq_85%2Cusm_0.66_1.00_0.01%2Cenc_auto/f4e552_94289803042f43d68a85e7c490b1fa1c~mv2.jpg?w=750&#038;ssl=1" alt=""/><figcaption class="wp-element-caption"><em>Digital art created using the A.I. DALL</em>·<em>E</em></figcaption></figure>



<p class="wp-block-paragraph" id="viewer-at41b"></p>



<p class="wp-block-paragraph" id="viewer-aqs8s">If a pipe is leaking from two holes and its pressure is kept fixed, then repairing one hole will result in the other one leaking faster. Similarly, as best practices increasingly become commonplace as a means to reduce p-hacking, so long as the career pressures to publish in top journals don&#8217;t let up, the occurrence of Importance Hacking may increase.</p>



<p class="wp-block-paragraph" id="viewer-3rjml">It&#8217;s time to start the conversation about how Importance Hacking can be addressed.</p>



<p class="wp-block-paragraph" id="viewer-agpq6">If you&#8217;re interested in learning more about Importance Hacking, you can listen to <a rel="noreferrer noopener" href="https://clearerthinkingpodcast.com/episode/122" target="_blank"><u>psychology professor Alexa Tullett and me discussing it on the Clearer Thinking podcast</u></a> (there, I refer to it as &#8220;Importance Laundering,&#8221; but I now think &#8220;Importance Hacking&#8221; is a better name) or me talking about it on the <a rel="noreferrer noopener" href="https://www.fourbeers.com/98" target="_blank"><u>Two Psychologists Four Beers podcast</u></a>. We also discuss my new project, <a rel="noreferrer noopener" href="https://replications.clearerthinking.org/" target="_blank"><u>Transparent Replications</u></a>, which conducts rapid replications of recently published psychology papers in top journals in an effort to shift incentives and create more reliable, replicable research. If you enjoyed this article, you may be interested in checking our <a rel="noreferrer noopener" href="https://replications.clearerthinking.org/replications/" target="_blank"><u>replication reports</u></a> and learning more <a rel="noreferrer noopener" href="https://replications.clearerthinking.org/about/" target="_blank"><u>about the project</u></a>.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph" id="viewer-es1me"><em>Did you like this article? If so, you may like to explore the ClearerThinking Podcast, where I have fun, in-depth conversations with brilliant people about ideas that matter. </em><a rel="noreferrer noopener" href="https://clearerthinkingpodcast.com/" target="_blank"><em><u>Click here to see a full list of episodes</u></em></a><em>.</em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2022/12/importance-hacking-a-major-yet-rarely-discussed-problem-in-science/feed/</wfw:commentRss>
			<slash:comments>6</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">3057</post-id>	</item>
		<item>
		<title>50 &#8220;Laws&#8221; of Everything</title>
		<link>https://www.spencergreenberg.com/2020/07/50-laws-of-everything-2/</link>
					<comments>https://www.spencergreenberg.com/2020/07/50-laws-of-everything-2/#respond</comments>
		
		<dc:creator><![CDATA[Spencer]]></dc:creator>
		<pubDate>Mon, 06 Jul 2020 23:06:00 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[adaptation]]></category>
		<category><![CDATA[attribution]]></category>
		<category><![CDATA[authority]]></category>
		<category><![CDATA[averages]]></category>
		<category><![CDATA[bureaucracy]]></category>
		<category><![CDATA[cognition]]></category>
		<category><![CDATA[communication]]></category>
		<category><![CDATA[complexity]]></category>
		<category><![CDATA[computation]]></category>
		<category><![CDATA[coordination]]></category>
		<category><![CDATA[corruption]]></category>
		<category><![CDATA[cosmology]]></category>
		<category><![CDATA[deadlines]]></category>
		<category><![CDATA[decision-making]]></category>
		<category><![CDATA[digital-rights]]></category>
		<category><![CDATA[drug-discovery]]></category>
		<category><![CDATA[economics]]></category>
		<category><![CDATA[efficiency]]></category>
		<category><![CDATA[emergence]]></category>
		<category><![CDATA[empathy]]></category>
		<category><![CDATA[estimation]]></category>
		<category><![CDATA[ethics]]></category>
		<category><![CDATA[evidence]]></category>
		<category><![CDATA[expectations]]></category>
		<category><![CDATA[experimentation]]></category>
		<category><![CDATA[expertise]]></category>
		<category><![CDATA[expertise-bias]]></category>
		<category><![CDATA[exponential-growth]]></category>
		<category><![CDATA[failure]]></category>
		<category><![CDATA[feedback-loops]]></category>
		<category><![CDATA[feminism]]></category>
		<category><![CDATA[forecasting]]></category>
		<category><![CDATA[forecasting-errors]]></category>
		<category><![CDATA[governance]]></category>
		<category><![CDATA[hierarchy]]></category>
		<category><![CDATA[human-nature]]></category>
		<category><![CDATA[ignorance]]></category>
		<category><![CDATA[incentives]]></category>
		<category><![CDATA[incompetence]]></category>
		<category><![CDATA[innovation]]></category>
		<category><![CDATA[institutions]]></category>
		<category><![CDATA[intelligence]]></category>
		<category><![CDATA[internet-culture]]></category>
		<category><![CDATA[language]]></category>
		<category><![CDATA[learning]]></category>
		<category><![CDATA[long-term-thinking]]></category>
		<category><![CDATA[management]]></category>
		<category><![CDATA[measurement]]></category>
		<category><![CDATA[media-literacy]]></category>
		<category><![CDATA[metrics]]></category>
		<category><![CDATA[misinformation]]></category>
		<category><![CDATA[morality]]></category>
		<category><![CDATA[network-effects]]></category>
		<category><![CDATA[neuroscience]]></category>
		<category><![CDATA[online-discussion]]></category>
		<category><![CDATA[optimization]]></category>
		<category><![CDATA[organizational-behavior]]></category>
		<category><![CDATA[ownership]]></category>
		<category><![CDATA[performance]]></category>
		<category><![CDATA[planning]]></category>
		<category><![CDATA[politics]]></category>
		<category><![CDATA[power]]></category>
		<category><![CDATA[prediction]]></category>
		<category><![CDATA[probability]]></category>
		<category><![CDATA[procrastination]]></category>
		<category><![CDATA[psychology]]></category>
		<category><![CDATA[quality]]></category>
		<category><![CDATA[randomness]]></category>
		<category><![CDATA[reform]]></category>
		<category><![CDATA[relationships]]></category>
		<category><![CDATA[scaling]]></category>
		<category><![CDATA[scientific-progress]]></category>
		<category><![CDATA[security]]></category>
		<category><![CDATA[skepticism]]></category>
		<category><![CDATA[social-dynamics]]></category>
		<category><![CDATA[social-networks]]></category>
		<category><![CDATA[software-development]]></category>
		<category><![CDATA[statistical-patterns]]></category>
		<category><![CDATA[statistics]]></category>
		<category><![CDATA[systems-thinking]]></category>
		<category><![CDATA[targets]]></category>
		<category><![CDATA[technological-change]]></category>
		<category><![CDATA[technology]]></category>
		<category><![CDATA[time-management]]></category>
		<category><![CDATA[tradeoffs]]></category>
		<category><![CDATA[tradition]]></category>
		<category><![CDATA[uncertainty]]></category>
		<category><![CDATA[unintended-consequences]]></category>
		<category><![CDATA[usefulness]]></category>
		<guid isPermaLink="false">https://www.spencergreenberg.com/?p=4892</guid>

					<description><![CDATA[This piece was first written on July 6, 2020, and first appeared on my website on May 30, 2026.]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph"></p>



<ol start="1" class="wp-block-list">
<li><strong>Parkinson&#8217;s Law</strong>: Work expands so as to fill the time available for its completion.</li>



<li><strong>Hofstadter&#8217;s Law</strong>: It always takes longer than you expect, even when you take into account Hofstadter’s Law.</li>



<li><strong>Gates&#8217; Law</strong>: Most people overestimate what they can do in one year and underestimate what they can do in ten years.</li>



<li><strong>Goodhart&#8217;s Law</strong>: When a measure becomes a target, it ceases to be a good measure.</li>



<li><strong>Hanlon&#8217;s Razor</strong>: Never attribute to malice that which is adequately explained by stupidity (or, don&#8217;t invoke conspiracy when ignorance and incompetence will suffice, as conspiracy implies intelligence).</li>



<li><strong>Acton&#8217;s Dictum</strong>: Power tends to corrupt, and absolute power corrupts absolutely.</li>



<li><strong>Amara&#8217;s Law</strong>: We tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run.</li>



<li><strong>Benford&#8217;s Law</strong>: In a diverse collection of unrelated statistics, a given statistic has roughly a 30% chance of starting with the digit 1.</li>



<li><strong>Betteridge&#8217;s Law</strong>: Any headline which ends in a question mark can be answered by the word &#8216;no&#8217;.</li>



<li><strong>Brooks&#8217; Law</strong>: Adding manpower to a late software project makes it later.</li>



<li><strong>Chesterton&#8217;s Fence</strong>: Reforms should not be made until the reasoning behind the existing state of affairs is understood.</li>



<li><strong>Claasen&#8217;s Law</strong>: Usefulness = log(technology).</li>



<li><strong>Clarke&#8217;s First Law</strong>: When a distinguished elderly scientist states that something is possible, they are almost certainly right, but when they state something is impossible, they are probably wrong.</li>



<li><strong>Cromwell&#8217;s Rule</strong>: Nothing but logical impossibilities have a prior probability of 0 or 1.</li>



<li><strong>Cunningham&#8217;s Law</strong>: The best way to get the right answer on the Internet is not to ask a question, it’s to post the wrong answer.</li>



<li><strong>Doctorow&#8217;s Law</strong>: When someone puts a lock on a thing you own, against your wishes, and doesn&#8217;t give you the key, they&#8217;re not doing it for your benefit.</li>



<li><strong>Dunbar&#8217;s Number</strong>: Most people can&#8217;t maintain stable social relationships with more than 150 people.</li>



<li><strong>Eroom&#8217;s Law</strong>: Drug discovery is becoming slower and more expensive over time, despite improvements in technology.</li>



<li><strong>Gell-Mann Amnesia Effect</strong>: You&#8217;ll believe articles outside your area of expertise, even after acknowledging that neighboring articles in your area of expertise are completely wrong.</li>



<li><strong>Gibson&#8217;s Law</strong> (or the Expert Witness Law): For each PhD (to use as an expert witness for one side) there&#8217;s an equal and opposite PhD.</li>



<li><strong>Godwin&#8217;s Law</strong>: As an online discussion grows longer, the probability of a comparison involving Nazis or Hitler approaches one.</li>



<li><strong>Morley-Souter&#8217;s Law</strong> (Rule 34): There is porn of it (no exceptions).</li>



<li><strong>Greenspun&#8217;s Tenth Rule</strong>: Any sufficiently complicated C program contains an ad hoc, informally specified, bug-ridden, slow implementation of half of Common Lisp.</li>



<li><strong>Hebb&#8217;s Law</strong>: Neurons that fire together wire together.</li>



<li><strong>Hubble&#8217;s Law</strong>: Galaxies recede from an observer at a rate proportional to their distance to that observer.</li>



<li><strong>Hume&#8217;s Guillotine</strong> (Is-Ought Problem): Normative statements (about what&#8217;s moral/immoral/right/wrong) cannot be deduced exclusively from descriptive statements.</li>



<li><strong>Humphrey&#8217;s Law</strong>: Conscious attention to a task normally performed automatically can impair its performance.</li>



<li><strong>Kranzberg&#8217;s Law</strong>: Technology is neither good nor bad; nor is it neutral.</li>



<li><strong>Lamarck&#8217;s Principle</strong> (or &#8220;Use it or Lose it&#8221;): Use it or lose it (evolutionarily speaking, but also in the brain).</li>



<li><strong>Lewis&#8217;s Law</strong>: The comments you&#8217;ll inevitably find on any article about feminism justify feminism.</li>



<li><strong>Littlewood&#8217;s Law</strong>: Individuals can expect miracles to happen to them, at the rate of about one per month.</li>



<li><strong>Maes–Garreau Law</strong>: Favorable predictions about future technology will fall at the latest possible date they can come true and still remain in the lifetime of the predictor.</li>



<li><strong>Metcalfe&#8217;s Law</strong>: The value of a system grows as approximately the square of the number of users of the system.</li>



<li><strong>Miller&#8217;s Law</strong>: To understand what another person is saying, you must assume that it is true and try to imagine what it could be true of.</li>



<li><strong>Moore&#8217;s Law</strong>: Computation per dollar grows exponentially (or: number of transistors per circuit doubles roughly every 24 months).</li>



<li><strong>Murphy&#8217;s Law</strong>: Anything that can go wrong will go wrong.</li>



<li><strong>Alder&#8217;s Law</strong>: What cannot be settled by experiment is not worth debating.</li>



<li><strong>O&#8217;Sullivan&#8217;s Law</strong>: All organizations that are not actually right-wing will over time become left-wing.</li>



<li><strong>Pareto&#8217;s Principle</strong> (80/20 Rule): For many phenomena 80% of consequences stem from 20% of the causes.</li>



<li><strong>Peter&#8217;s Principle</strong>: In a hierarchy, every employee tends to rise to his level of incompetence.</li>



<li><strong>Poisson&#8217;s Law</strong> (or Law of Large Numbers): For independent random variables with a common distribution, the average tends to the mean as sample size increases.</li>



<li><strong>Pournelle&#8217;s Iron Law of Bureaucracy</strong>: In bureaucracy, those devoted to the bureaucracy get control, those devoted to what it&#8217;s supposed to achieve lose influence.</li>



<li><strong>Putt&#8217;s Law</strong>: Technology is dominated by two types of people: those who understand what they do not manage and those who manage what they do not understand.</li>



<li><strong>Rosenthal Effect</strong> (Pygmalion Effect): High expectations lead to an increase in performance, low expectations to a decrease in performance.</li>



<li><strong>Schneier&#8217;s Law</strong>: Any person can invent a security system so clever that she or he can&#8217;t think of how to break it.</li>



<li><strong>Shermer&#8217;s Law</strong>: Any sufficiently advanced extraterrestrial intelligence is indistinguishable from God.</li>



<li><strong>Zipf&#8217;s Law</strong>: The frequency of use of the nth-most-frequently-used word in any natural language is approximately inversely proportional to n (few words are used often, most are used rarely).</li>



<li><strong>Wirth&#8217;s Law</strong>: Software gets slower more quickly than hardware gets faster.</li>



<li><strong>Sturgeon&#8217;s Law</strong>: Ninety percent of everything is crud.</li>



<li><strong>Stigler&#8217;s Law</strong>: No discovery is named after its original discoverer, including this one.</li>
</ol>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>This piece was first written on July 6, 2020, and first appeared on my website on May 30, 2026.</em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2020/07/50-laws-of-everything-2/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">4892</post-id>	</item>
		<item>
		<title>The many ways to make inferences</title>
		<link>https://www.spencergreenberg.com/2018/10/the-many-ways-to-make-inferences/</link>
					<comments>https://www.spencergreenberg.com/2018/10/the-many-ways-to-make-inferences/#comments</comments>
		
		<dc:creator><![CDATA[admin]]></dc:creator>
		<pubDate>Mon, 08 Oct 2018 01:44:00 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[Bayesianism]]></category>
		<category><![CDATA[cases]]></category>
		<category><![CDATA[causal analysis]]></category>
		<category><![CDATA[causes]]></category>
		<category><![CDATA[deduction]]></category>
		<category><![CDATA[experts]]></category>
		<category><![CDATA[frequentism]]></category>
		<category><![CDATA[inference]]></category>
		<category><![CDATA[intuition]]></category>
		<category><![CDATA[metaphors]]></category>
		<category><![CDATA[modeling]]></category>
		<category><![CDATA[probabilistic reasoning]]></category>
		<category><![CDATA[probability]]></category>
		<category><![CDATA[rationality]]></category>
		<category><![CDATA[reasoning]]></category>
		<category><![CDATA[regression]]></category>
		<category><![CDATA[similarities]]></category>
		<category><![CDATA[theorizing]]></category>
		<guid isPermaLink="false">https://www.spencergreenberg.com/?p=2564</guid>

					<description><![CDATA[There are a LOT of ways to make inferences. Many more, I think, than is generally realized. And they all have their weaknesses. You can make inferences using… (1) Deduction: As a consequence of the definition of X and Y, if X then Y. X applies to this case. Therefore Y. “Plato is a man, [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">There are a LOT of ways to make inferences. Many more, I think, than is generally realized. And they all have their weaknesses.</p>



<p class="wp-block-paragraph">You can make inferences using…</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph"><strong>(1) Deduction:</strong></p>



<p class="wp-block-paragraph">As a consequence of the definition of X and Y, if X then Y.</p>



<p class="wp-block-paragraph">X applies to this case. Therefore Y.</p>



<p class="wp-block-paragraph">“Plato is a man, and all men are mortal; therefore Plato is mortal.”</p>



<p class="wp-block-paragraph">“For any number that is an integer, there exists another integer greater than that number. 1,000,000 is an integer. So there exists an integer greater than 1,000,000.”</p>



<p class="wp-block-paragraph"><em>Especially popular among </em>philosophers and mathematicians?</p>



<p class="wp-block-paragraph"><em>Flaws:</em> to apply to the world, you need to add in assumptions about the world, or to apply other methods of inference on top.</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph"><strong>(2) Frequencies:</strong></p>



<p class="wp-block-paragraph">In the past, 95% of the time that X occurred, Y occurred.</p>



<p class="wp-block-paragraph">X occurred. Therefore Y (with high probability).</p>



<p class="wp-block-paragraph">“95% of the time when we saw a transaction identical to this one, it was fraudulent. So this transaction is fraudulent.”</p>



<p class="wp-block-paragraph"><em>Especially popular among </em>applied statisticians and data scientists?</p>



<p class="wp-block-paragraph"><em>Flaws: </em>You need to have a moderately large number of examples like the current one to perform calculations on, and the method assumes that those past examples were drawn from a process that is (statistically) just like the one that generated this latest example. Moreover, sometimes it is unclear what it means for “X” to have occurred. What if it’s something that’s very similar to but not quite like X that occurred &#8211; should that be counted? If we broaden our class of what counts or change to another class that still encompasses all of our prior examples, we’ll potentially get a different answer. Though, fortunately, there are plenty of cases where the class to use is fairly obvious.</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph"><strong>(3) Models:</strong></p>



<p class="wp-block-paragraph">Given our probabilistic model of this thing, when X occurs, the probability of Y is 0.95.</p>



<p class="wp-block-paragraph">X occurred. Therefore Y (with high probability).</p>



<p class="wp-block-paragraph">“Given our multivariate Gaussian model of loan prices, when this loan defaults, there is a 0.95 probability of this other loan defaulting.”</p>



<p class="wp-block-paragraph"><em>Especially popular among</em> financial engineers and risk modelers?</p>



<p class="wp-block-paragraph"><em>Flaws:</em> hinges on the appropriateness of the parameterized probabilistic model chosen, may require a moderately large amount of past data to estimate free model parameters, and may go haywire if modeling assumptions are suddenly violated.</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph"><strong>(4) Regression:</strong></p>



<p class="wp-block-paragraph">In prior data, as X and Z increased, the likelihood of Y increased.</p>



<p class="wp-block-paragraph">X and Z are at high levels. Therefore Y.</p>



<p class="wp-block-paragraph">“Height for children can be approximately predicted as an (increasing/positive) linear function of age and weight. This child is older and heavier than the others, so we predict he is also taller than the others.”</p>



<p class="wp-block-paragraph"><em>Especially common among </em>economists and data scientists?</p>



<p class="wp-block-paragraph"><em>Flaws:</em> often is applied with simple assumptions (e.g., linearity) that may not capture the complexity of the inference, but very large amounts of data may be needed to apply much more complex models (e.g., to use neural networks).</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph"><strong>(5) Bayesianism:</strong></p>



<p class="wp-block-paragraph">Given my prior odds on Y being true…</p>



<p class="wp-block-paragraph">And given evidence X…</p>



<p class="wp-block-paragraph">And given my Bayes factor, which is my estimate of how much more likely X is to occur if Y is true than if Y is not true…</p>



<p class="wp-block-paragraph">I calculate that Y is far more likely to be true than to not be true (by multiplying the prior odds by the Bayes factor to get the posterior odds).</p>



<p class="wp-block-paragraph">Therefore Y (with high probability).</p>



<p class="wp-block-paragraph">“My prior odds that my boss is angry at me were 1 to 4, because he’s angry at me about 20% of the time. But then he came into my office shouting and flipped over my desk, which I estimate is 200 times more likely to occur if he’s angry at me compared to if he’s not. So now the odds of him being angry at me are 200 * (1/4) = 50 to 1 in favor of him being angry.”</p>



<p class="wp-block-paragraph"><em>Not as popular as it should be?</em></p>



<p class="wp-block-paragraph"><em>Flaws:</em> it is sometimes hard to know what to set your prior odds at, and it can be very hard in some cases to perform the calculation. In practice, carrying out the calculation might end up relying on subjective estimates of the odds, which can be especially tricky to guess when the evidence is not binary (i.e., not of the form “happened” vs. “didn’t happened”), or if you have lots of different pieces of evidence that are partially correlated. On the other hand, if you can do the calculations in a given instance, and have a sensible way to set a prior, this is, in my opinion, the mathematically optimal framework to use for probabilistic prediction. In that sense, we can think of many of the other approaches on this list as (hopefully pragmatic) approximations of Bayesianism (sometimes good approximations, sometimes bad ones).</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph">(<strong>6) Theories:</strong></p>



<p class="wp-block-paragraph">Given our theory, when X occurs, Y occurs.</p>



<p class="wp-block-paragraph">X occurred. Therefore Y.</p>



<p class="wp-block-paragraph">“One theory is that depressed people are most at risk for suicide when they are beginning to come out of a really bad depressive episode. So as depression is remitting, patients should be carefully screened for potentially increasing risk factors.”</p>



<p class="wp-block-paragraph">“When inflation rises, unemployment falls. Inflation is rising, so unemployment will fall.”</p>



<p class="wp-block-paragraph"><em>Especially popular among </em>psychologists and economists?</p>



<p class="wp-block-paragraph"><em>Flaws:</em> it’s very challenging to come up with reliable theories, and often you will not know how accurate such a theory is. Even if it has substantial truth to it and is often right, there may be cases where the opposite of what was predicted actually happens, and for reasons that the theory can’t explain.</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph"><strong>(7) Causes:</strong></p>



<p class="wp-block-paragraph">We know that X causes Y to occur.</p>



<p class="wp-block-paragraph">X occurred. Therefore Y.</p>



<p class="wp-block-paragraph">“Rusting of gears causes increased friction, leading to greater wearing down. In this case, the gears were heavily rusted, so we expect to find a lot of wearing down.”</p>



<p class="wp-block-paragraph">“This gene produces this phenotype, and we see that this gene is present, so we expect to see the phenotype.”</p>



<p class="wp-block-paragraph"><em>Especially popular among</em> engineers and biologists?</p>



<p class="wp-block-paragraph"><em>Flaws: </em>it’s often extremely hard to figure out causality in a highly complex system, especially in “softer” subjects like nutrition.</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph"><strong>(8)</strong> <strong>Experts:</strong></p>



<p class="wp-block-paragraph">This expert (or prediction market, or prediction algorithm) X is 90% accurate at predicting things in this general domain of prediction.</p>



<p class="wp-block-paragraph">X predicts Y. Therefore Y (with high probability).</p>



<p class="wp-block-paragraph">“This prediction market has been right 90% of the time when predicting recent baseball outcomes, and in this case, they predict that the Yankees will win.”</p>



<p class="wp-block-paragraph"><em>Not as popular as it should be?</em></p>



<p class="wp-block-paragraph"><em>Flaws: </em>you often don’t have access to the predictions of experts (or of prediction markets or prediction algorithms), and when you do, you usually don’t have reliable measures of their past accuracy.</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph"><strong>(9) Metaphors:</strong></p>



<p class="wp-block-paragraph">X, which is what we are dealing with now, is metaphorically a Z.</p>



<p class="wp-block-paragraph">For Z, when W is true, then obviously Y.</p>



<p class="wp-block-paragraph">Now W (or its metaphorical equivalent) is true for X. Therefore Y.</p>



<p class="wp-block-paragraph">“Your life is but a boat, and you are riding on the waves of your experiences. When a raging storm hits, a boat can’t be under full sail. It can’t continue at its maximum speed. You are experiencing a storm now, and so you too must learn to slow down.”</p>



<p class="wp-block-paragraph"><em>Especially popular among</em> self-help gurus and some ancient philosophers?</p>



<p class="wp-block-paragraph"><em>Flaws:</em> Z working as a metaphor for X doesn’t mean that all (or even most) solutions that are good for situations involving Z are appropriate (or even make any sense) for X.</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph"><strong>(10) Similarities:</strong></p>



<p class="wp-block-paragraph">X occurred, and X is very similar to Z in properties A, B, and C.</p>



<p class="wp-block-paragraph">When things similar to Z in properties A, B, and C occur, Y usually occurs.</p>



<p class="wp-block-paragraph">Therefore Y (with high probability).</p>



<p class="wp-block-paragraph">“This conflict is similar to the Gulf war in that&#8230;and with “Gulf”-like wars, we have always seen that&#8230;”</p>



<p class="wp-block-paragraph">“This data point (with unknown label) is closest in feature space to this other data point which is labeled ‘cat,’ and all the other labeled points around that point are also labeled ‘cat,’ so this unlabeled point should also likely get the label ‘cat.’”</p>



<p class="wp-block-paragraph"><em>Especially popular among</em> historians and within some machine learning algorithms?</p>



<p class="wp-block-paragraph"><em>Flaws:</em> in the history case, it is difficult to know which features are the appropriate ones to use to compare similarities, and often the conclusions are based on a relatively small number of examples. In the machine learning case, a very large amount of data may be needed to train the model.</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph"><strong>(11) Cases:</strong></p>



<p class="wp-block-paragraph">In this handful of examples (or perhaps even just one example) where X occurred, Y occurred.</p>



<p class="wp-block-paragraph">X occurred. Therefore Y.</p>



<p class="wp-block-paragraph">“The last time we elected a [insert political group you don’t like] as president, we saw what happened. Let’s not make that mistake again.”</p>



<p class="wp-block-paragraph">“The last three times I went to action movies, I didn’t like them. So I don’t want to go to one again.”</p>



<p class="wp-block-paragraph"><em>Especially popular with</em> politicians and with nearly everyone in daily living?</p>



<p class="wp-block-paragraph"><em>Flaws:</em> unless we are in a situation with very little noise/variability, a few examples likely will not be enough to accurately generalize from.</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph"><strong>(12) Intuition:</strong></p>



<p class="wp-block-paragraph">X occurred. My intuition (that I may have trouble explaining) predicts that when X occurs, Y is true. Therefore Y.</p>



<p class="wp-block-paragraph">“The tone of voice he used when he talked about his family gave me a bad vibe. My feeling is that anyone who talks about their family with that tone of voice probably does not really love them.”</p>



<p class="wp-block-paragraph"><em>Popular with</em> nearly everyone in daily living?</p>



<p class="wp-block-paragraph"><em>Flaws: </em>our intuitions can be very well-honed in situations we’ve encountered many times and that we received feedback on (i.e., where there was some sort of answer we got about how well our intuition performed), but in highly novel situations or in situations where we receive no feedback on how well our intuition is performing, our intuitions may be highly inaccurate (even though we may not FEEL any less confident about our correctness).</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph"><em>This essay was first written on October 7, 2018, and first appeared on this site on December 31, 2021.</em>&nbsp;</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2018/10/the-many-ways-to-make-inferences/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">2564</post-id>	</item>
		<item>
		<title>A thought experiment about what you&#8217;d be truly capable of doing (if you had no choice)</title>
		<link>https://www.spencergreenberg.com/2018/04/a-thought-experiment-about-what-youd-be-truly-capable-of-doing-if-you-had-no-choice/</link>
					<comments>https://www.spencergreenberg.com/2018/04/a-thought-experiment-about-what-youd-be-truly-capable-of-doing-if-you-had-no-choice/#respond</comments>
		
		<dc:creator><![CDATA[admin]]></dc:creator>
		<pubDate>Thu, 26 Apr 2018 20:33:00 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[existential risks]]></category>
		<category><![CDATA[focus]]></category>
		<category><![CDATA[hypothetical]]></category>
		<category><![CDATA[possibility]]></category>
		<category><![CDATA[probability]]></category>
		<category><![CDATA[self-efficacy]]></category>
		<guid isPermaLink="false">https://www.spencergreenberg.com/?p=2455</guid>

					<description><![CDATA[Think of something you value that:A. multiple other people you know are capable of achieving, but thatB. you assume you would not be capable of achieving, even thoughC. you have never actually tried to do this thing well before. Now suppose, for a moment, that you have no choice but to do the thing. That [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Think of something you value that:<br>A. multiple other people you know are capable of achieving, but that<br>B. you assume you would not be capable of achieving, even though<br>C. you have never actually tried to do this thing well before.</p>



<p class="wp-block-paragraph">Now suppose, for a moment, that you have no choice but to do the thing. That is, everything you care about in the world will be destroyed if you do not achieve it in X months. Here, X could be 1 if it&#8217;s a very small thing, or X could be 100 if it&#8217;s a much larger thing.</p>



<p class="wp-block-paragraph">Under those circumstances, do you STILL believe you would fail to achieve it?</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph">I think this sort of thought experiment can help us distinguish between things that we don&#8217;t believe we are capable of merely because we aren&#8217;t motivated enough versus things that we ACTUALLY believe are impossible for us.</p>



<p class="wp-block-paragraph">And I think it&#8217;s important to distinguish between these two cases, because if something is in the first category, we may actually be able to get ourselves to succeed just by finding ways to increase our motivation!</p>



<hr class="wp-block-separator"/>



<p class="wp-block-paragraph">I also suspect that for many people, a number of the things that they view as being impossible for them would be more likely to seem possible in the face of carrying out this thought experiment. In other words, it is easy to confuse &#8220;I&#8217;m not motivated enough to try really hard&#8221; with &#8220;I&#8217;m incapable.&#8221;</p>



<p class="wp-block-paragraph">As an example: suppose that you believe you are just inherently bad at math and that no matter how hard you try, you couldn&#8217;t understand calculus. Well, what if the fate of the world rested on you understanding calculus in six months? Under those circumstances, I think you would very likely find a way to learn it, with plenty of time to spare.</p>



<p class="wp-block-paragraph"><em>This piece was first written on April 26, 2018, and was first released on this site on October 1, 2021.</em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2018/04/a-thought-experiment-about-what-youd-be-truly-capable-of-doing-if-you-had-no-choice/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">2455</post-id>	</item>
		<item>
		<title>Which Risks of Dying Are Worth Taking?</title>
		<link>https://www.spencergreenberg.com/2013/10/which-risks-of-dying-are-worth-taking/</link>
					<comments>https://www.spencergreenberg.com/2013/10/which-risks-of-dying-are-worth-taking/#comments</comments>
		
		<dc:creator><![CDATA[Spencer]]></dc:creator>
		<pubDate>Wed, 16 Oct 2013 19:20:47 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[chance]]></category>
		<category><![CDATA[danger]]></category>
		<category><![CDATA[death]]></category>
		<category><![CDATA[luck]]></category>
		<category><![CDATA[probability]]></category>
		<category><![CDATA[risk]]></category>
		<guid isPermaLink="false">http://www.spencergreenberg.com/?p=777</guid>

					<description><![CDATA[First, click here to figure out your chance of dying tomorrow. Is it worth taking a 1 in 100,000 chance of dying , in order to experience the novel thrill of sky diving? Is a 1 in 500,000 chance of death worth it to go bungee jumping? It&#8217;s hard to know whether these risks are [&#8230;]]]></description>
										<content:encoded><![CDATA[<h3><a href="http://gimbeltech.com/clearerthinking/death.html" target="_blank">First, click here to figure out your chance of dying tomorrow.</a></h3>
<p>Is it worth taking a 1 in 100,000 chance of dying , in order to experience the novel thrill of sky diving? Is a 1 in 500,000 chance of death worth it to go bungee jumping?</p>
<p>It&#8217;s hard to know whether these risks are reasonable, because numbers like 100,000 or 500,000 feel so abstract to us. To think more clearly about these numbers, it helps to get our intuitions engaged.</p>
<p>We can start by figuring out the daily risk of dying that we automatically face every day. My &#8220;death calculator&#8221; tool above will compute yours, as estimated from your gender and age. For instance, in the United States, a 30 year old man has about a 1 in 260,000 chance of dying tomorrow whereas a 30 year old woman has about a 1 in 583,000 chance. A 55 year old man has a 1 in 46,000 chance of dying on any given day and a 55 year old woman a 1 in 79,000 chance. Note that while it&#8217;s extremely difficult to estimate a person&#8217;s life span (since future technological and societal changes may radically alter how long people live), estimating how likely a person is to die in the next day is much more accurate and straightforward.</p>
<p>Once you&#8217;ve used the tool to calculate your own chance of dying tomorrow, you can start thinking about the risk of dangerous activities relative to how much risk you already take each day (merely by going about normal activities). In particular, you can calculate how many total days worth of risk an activity involves. So for instance, if you are a 30 year old male, and ride 100 miles on a motorcycle tomorrow, then you&#8217;ll experienced 11.2 days worth of risk of dying tomorrow, rather than a single normal day of risk.</p>
<p>To do the calculation of how many days of risk you&#8217;re taking in a day where you do the dangerous activity, simply calculate the following: Start with the probability that you die in a normal day, add to it the probability that you die from doing the risky activity, and then divide the result by the probability that you die in a normal day.</p>
<p>So for instance, if you were to go BASE jumping tomorrow (an activity that appears to have about a 1 in 2,300 chance of death), and if you normally have a 1 in 100,000 chance of dying in a given day (for instance, you&#8217;re a 46 year old man) then you&#8217;d be taking on ((1/2300)+(1/100000))/(1/100,000) = 44.5 days worth of ordinary daily risk tomorrow, instead of just 1 day of risk.</p>
<p><a href="https://i0.wp.com/www.spencergreenberg.com/wp-content/uploads/2013/10/BASE_Jumping_from_Sapphire_Tower_in_Istanbul.jpg"><img data-recalc-dims="1" fetchpriority="high" decoding="async" data-attachment-id="804" data-permalink="https://www.spencergreenberg.com/2013/10/which-risks-of-dying-are-worth-taking/base_jumping_from_sapphire_tower_in_istanbul/" data-orig-file="https://i0.wp.com/www.spencergreenberg.com/wp-content/uploads/2013/10/BASE_Jumping_from_Sapphire_Tower_in_Istanbul.jpg?fit=4288%2C2848&amp;ssl=1" data-orig-size="4288,2848" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;8&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;NIKON D5000&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;1337335408&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;18&quot;,&quot;iso&quot;:&quot;400&quot;,&quot;shutter_speed&quot;:&quot;0.0005&quot;,&quot;title&quot;:&quot;&quot;}" data-image-title="BASE_Jumping_from_Sapphire_Tower_in_Istanbul" data-image-description="" data-image-caption="" data-large-file="https://i0.wp.com/www.spencergreenberg.com/wp-content/uploads/2013/10/BASE_Jumping_from_Sapphire_Tower_in_Istanbul.jpg?fit=750%2C498&amp;ssl=1" class="aligncenter size-full wp-image-804" alt="BASE_Jumping_from_Sapphire_Tower_in_Istanbul" src="https://i0.wp.com/www.spencergreenberg.com/wp-content/uploads/2013/10/BASE_Jumping_from_Sapphire_Tower_in_Istanbul.jpg?resize=750%2C498" width="750" height="498" srcset="https://i0.wp.com/www.spencergreenberg.com/wp-content/uploads/2013/10/BASE_Jumping_from_Sapphire_Tower_in_Istanbul.jpg?w=4288&amp;ssl=1 4288w, https://i0.wp.com/www.spencergreenberg.com/wp-content/uploads/2013/10/BASE_Jumping_from_Sapphire_Tower_in_Istanbul.jpg?resize=300%2C199&amp;ssl=1 300w, https://i0.wp.com/www.spencergreenberg.com/wp-content/uploads/2013/10/BASE_Jumping_from_Sapphire_Tower_in_Istanbul.jpg?resize=1024%2C680&amp;ssl=1 1024w, https://i0.wp.com/www.spencergreenberg.com/wp-content/uploads/2013/10/BASE_Jumping_from_Sapphire_Tower_in_Istanbul.jpg?w=1500&amp;ssl=1 1500w, https://i0.wp.com/www.spencergreenberg.com/wp-content/uploads/2013/10/BASE_Jumping_from_Sapphire_Tower_in_Istanbul.jpg?w=2250&amp;ssl=1 2250w" sizes="(max-width: 750px) 100vw, 750px" /></a></p>
<p>Would that be worth it? For <em>some</em> people, it <em>might</em> be possible it is worth BASE jumping once in their in life. But suppose you were to go BASE jumping 20 times over the next year, on 20 different days. Then rather than consuming 365 days of typical risk that year (as a 46 year old man), you&#8217;d be taking on about 1235 days worth of risk, an additional roughly 2.4 years of risk! It&#8217;s hard to imagine that being worth it, even if BASE jumping is incredibly enjoyable. Of course, there is also a high risk of injury, aside from the risk of death.</p>
<p>Below is a table with estimates of the chance of dying from doing various activities.</p>
<p>Taking a 340 mile road trip on occasion with friends seems very reasonable. For instance, a 30 year old male will only be doubling his risk of dying that day, and a 30 year old female will be taking on about 3.3 days of her usual daily risk. But taking a job as a taxi driver in a suburban area or a long distance courier, driving 340 miles most days, would be much more risky. A 30 year old male who took such a job would be doubling his risk of dying <em>every</em> day. Another way to think about it is that despite being a 30 year old male, he would living with the daily risk of a 43 year old male. Similarly, a 30 year old male who decided to go BASE jumping one day, would be living that day with the daily risk of death of an 88 year old man.</p>
<p>So what risks are worth taking? It&#8217;s ultimately a subjective question. But thinking in terms of how much you&#8217;re increasing your ordinary daily risk, and converting risks into the daily risk of people of different ages, can make these abstract numbers more intuitive. And stronger intuition can help us reason more sanely about our choices.</p>
<h6><span style="color: #ff0000;"><strong>Disclaimer</strong></span>: these numbers come from the internet, so if you want to be confident about any of them, you should double check them using a reliable source. Note also that actual death rates will depend on a variety of factors, including the amount of experience you have doing that activity and the location where you are doing it. Furthermore, note that for many of these activities there are risks of injury, in addition to risk of death, but I&#8217;ve only considered risk of death in this analysis. Be very cautious before deciding to engage in any dangerous activity!</h6>
<table width="1417" border="0" cellspacing="0" cellpadding="0">
<colgroup>
<col width="321" />
<col width="202" />
<col width="135" />
<col width="251" />
<col width="400" />
<col width="260" /> </colgroup>
<tbody>
<tr>
<td width="321" height="15">
<address><strong>Activity</strong></address>
</td>
<td width="202">
<address><strong>1 in __ chance of dying</strong></address>
</td>
<td width="500">
<address><strong>per</strong></address>
</td>
<td width="350">
<address><strong>Days of risk (30 yr male)</strong></address>
</td>
<td width="350">
<address><strong>Days of risk (30 yr female)</strong></address>
</td>
<td width="260">
<address><strong> </strong></address>
</td>
</tr>
<tr>
<td height="15">
<address>BASE Jumping</address>
</td>
<td>
<address>2,300</address>
</td>
<td>
<address>jump</address>
</td>
<td>
<address>113.6</address>
</td>
<td>
<address>252.6</address>
</td>
<td>
<address><a href="http://www.medicine.ox.ac.uk/bandolier/booth/risk/sports.html">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Driving 100 miles on a motorcycle</address>
</td>
<td>
<address>26,000</address>
</td>
<td>
<address>100 miles driven</address>
</td>
<td>
<address>11.2</address>
</td>
<td>
<address>23.7</address>
</td>
<td>
<address><a href="http://www-nrd.nhtsa.dot.gov/Pubs/810990.PDF">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Participating in a triathlon</address>
</td>
<td>
<address>67,000</address>
</td>
<td>
<address>run</address>
</td>
<td>
<address>4.9</address>
</td>
<td>
<address>9.7</address>
</td>
<td>
<address><a href="http://well.blogs.nytimes.com/2009/10/20/are-marathons-safe/">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Skydiving</address>
</td>
<td>
<address>101,000</address>
</td>
<td>
<address>jump</address>
</td>
<td>
<address>3.6</address>
</td>
<td>
<address>6.8</address>
</td>
<td>
<address><a href="http://www.medicine.ox.ac.uk/bandolier/booth/risk/sports.html">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Hang-gliding</address>
</td>
<td>
<address>116,000</address>
</td>
<td>
<address>flight</address>
</td>
<td>
<address>3.2</address>
</td>
<td>
<address>6.0</address>
</td>
<td>
<address><a href="http://www.medicine.ox.ac.uk/bandolier/booth/risk/sports.html">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Running a marathon (cardiac arrest)</address>
</td>
<td>
<address>127,000</address>
</td>
<td>
<address>run</address>
</td>
<td>
<address>3.1</address>
</td>
<td>
<address>5.6</address>
</td>
<td>
<address><a href="http://www.medicine.ox.ac.uk/bandolier/booth/risk/sports.html">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Scuba Diving</address>
</td>
<td>
<address>200,000</address>
</td>
<td>
<address>dive</address>
</td>
<td>
<address>2.3</address>
</td>
<td>
<address>3.9</address>
</td>
<td>
<address><a href="http://www.medicine.ox.ac.uk/bandolier/booth/risk/sports.html">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Rock Climbing</address>
</td>
<td>
<address>320,000</address>
</td>
<td>
<address>climb</address>
</td>
<td>
<address>1.8</address>
</td>
<td>
<address>2.8</address>
</td>
<td>
<address><a href="http://www.medicine.ox.ac.uk/bandolier/booth/risk/sports.html">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Bungee Jumping</address>
</td>
<td>
<address>500,000</address>
</td>
<td>
<address>jump</address>
</td>
<td>
<address>1.5</address>
</td>
<td>
<address>2.2</address>
</td>
<td>
<address><a href="http://www.ehow.com/how_2053210_bungee-jump.html">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Canoeing</address>
</td>
<td>
<address>750,000</address>
</td>
<td>
<address>outing</address>
</td>
<td>
<address>1.3</address>
</td>
<td>
<address>1.8</address>
</td>
<td>
<address><a href="http://www.medicine.ox.ac.uk/bandolier/booth/risk/sports.html">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Driving 100 miles in a car</address>
</td>
<td>
<address>877,000</address>
</td>
<td>
<address>100 miles driven</address>
</td>
<td>
<address>1.3</address>
</td>
<td>
<address>1.7</address>
</td>
<td>
<address><a href="http://www.hcra.harvard.edu/quiz.html">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Skiing</address>
</td>
<td>
<address>1,557,000</address>
</td>
<td>
<address>visit</address>
</td>
<td>
<address>1.2</address>
</td>
<td>
<address>1.4</address>
</td>
<td>
<address><a href="http://www.medicine.ox.ac.uk/bandolier/booth/risk/sports.html">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Snow Boarding</address>
</td>
<td>
<address>2,198,000</address>
</td>
<td>
<address>visit</address>
</td>
<td>
<address>1.1</address>
</td>
<td>
<address>1.3</address>
</td>
<td>
<address><a href="http://www.ski-injury.com/prevention/helmet">source</a></address>
</td>
</tr>
<tr>
<td height="15">
<address>Flying 1,000 miles in a commercial airplane in the U.S.</address>
</td>
<td>
<address>3,333,000</address>
</td>
<td>
<address>1,000 miles flown</address>
</td>
<td>
<address>1.1</address>
</td>
<td>
<address>1.2</address>
</td>
<td>
<address><a href="http://en.wikipedia.org/wiki/Aviation_safety#United_States">source</a></address>
</td>
</tr>
</tbody>
</table>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2013/10/which-risks-of-dying-are-worth-taking/feed/</wfw:commentRss>
			<slash:comments>4</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">777</post-id>	</item>
		<item>
		<title>Testing Too Many Hypotheses</title>
		<link>https://www.spencergreenberg.com/2011/10/testing-too-many-hypotheses/</link>
					<comments>https://www.spencergreenberg.com/2011/10/testing-too-many-hypotheses/#comments</comments>
		
		<dc:creator><![CDATA[Spencer]]></dc:creator>
		<pubDate>Mon, 10 Oct 2011 17:16:40 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[experiments]]></category>
		<category><![CDATA[hypotheses]]></category>
		<category><![CDATA[hypothesis test]]></category>
		<category><![CDATA[p-values]]></category>
		<category><![CDATA[probability]]></category>
		<category><![CDATA[science]]></category>
		<category><![CDATA[statistics]]></category>
		<category><![CDATA[test]]></category>
		<guid isPermaLink="false">http://www.spencergreenberg.com/?p=249</guid>

					<description><![CDATA[For each dataset, there is a limit to what we can use that dataset to test. Using the standard p-value based methods of science, the more hypotheses we check against the data, the more likely it will be that some of these checks give inaccurate conclusions. And this presents a big problem for the way [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>For each dataset, there is a limit to what we can use that dataset to test. Using the standard <a href="http://en.wikipedia.org/wiki/P-value">p-value based methods</a> of science, the more hypotheses we check against the data, the more likely it will be that some of these checks give inaccurate conclusions. And this presents a big problem for the way science is practiced.</p>
<p>Let&#8217;s take an example to illustrate the principle. Suppose that you have information about 1000 people selected at random from the U.S. adult population. Your dataset includes these people&#8217;s heights, weights, ages, shoe sizes, and so forth. Now, if your goal is to know the mean height of all people in America, you can produce an estimate of this quantity by averaging the heights of the 1000 people you have information about. Despite the fact that your sample contains just 1000 people, rather than the full set of 230,000,000 or so American adults of interest, your estimate will, with high probability, be within a couple of inches of the total population mean height. This is due to the fact that the 1000 people were sampled at random (so we shouldn&#8217;t expect our sample to differ from the entire population in a systematic way) and because the standard deviation of heights is not very large (if there were tremendous outliers in the data, such as 500 foot tall giants, we would need more samples to get an accurate estimate). This idea is made precise by the central limit theorem. It tells us how likely the true entire population mean is to fall different distances from our sample estimate, and says that the error of our estimate decreases like one over the square root of the size of our sample.</p>
<p>The same technique could work to approximate the mean weight of adult Americans, or age, or shoe size, or number of children. And in each case, the estimate would, with high probability, be quite accurate. We could even, if we liked, estimate all these quantities simultaneously if we collected all of this information about each of our 1000 people. But the more quantities we estimate, the greater the chance that at least one estimate is quite inaccurate. Since each estimate has some chance of being bad, if we make a sufficiently large number of estimates we should expect to get unlucky at some point and end up with one or more bad ones. So, if we aren&#8217;t just estimating mean height, but rather the mean of 50 different traits, we cannot claim that all 50 of these estimate are likely to be good. We should expect that some of them will be inaccurate, though we don&#8217;t know which ones.</p>
<p>This is where problems arise. Suppose that you are a researcher who is trying to find interesting differences between, say, southerners and northerners in the United States. Your dataset of 1000 adults contains 500 people from each group. What do you do? Well, it might seem reasonable to go ahead and compute the mean value of many different traits, and look at how these means differ between the two groups, to see if you can find any large differences that seem interesting. For instance, you may compute the average salary of each group, and see if they deviate from each other by a large enough amount to be deemed statistically significant. If they don&#8217;t, you can try another trait like IQ, or number of children, and repeat the process. If you try enough different traits, hopefully you&#8217;ll eventually find an intriguingly large difference between the groups.</p>
<p>The trouble is, we know that if you estimate a large number of quantities, some of them will be inaccurate, and so some of the apparent differences between your two groups may just be due to these inaccuracies. If you test enough traits, you will eventually find differences between the populations that look significant, even though it is just the result of chance.</p>
<p>In fact, even if northerners and southerners had no systematic differences between them, there would still be apparent differences that arose just from the particular sample of 1000 people you happened to have data on. For example, in your dataset, it just might happen that the northerners have lower numbers of children than southerners, even if this isn&#8217;t true for the underlying populations of all northerners and southerners. If you were to publish this finding, without making mention of the number of hypotheses you tested before finding it, it may seem that you had produced a meaningful result. In fact, the assessment of this result should take into account the number of hypotheses (e.g. northerners have smaller shoe sizes than southerners, northerners have greater salaries than southerners, etc.) that you tested before you discovered this one (and the <a href="http://en.wikipedia.org/wiki/Multiple_comparisons">p-values can be modified to include this information</a>). The most significant seeming deviation between the groups found after testing 100 different hypotheses is very likely greatly inflated by chance. Whereas if you had only tested a small number of hypotheses against your data, and found a strong result, this would likely be a meaningful finding.</p>
<p>As a general rule, the greater the number of data points you have, the larger the number of quantities you can accurately estimate from your dataset. On a set of just 10 points, you may not even be able to get an accurate estimate of the mean value of a single trait (unless the trait had very slow standard deviation). Whereas on a dataset of a billion points, you probably could estimate dozens of quantities accurately.</p>
<p>Unfortunately, when you&#8217;re reading a paper, there is no way to tell how many hypotheses the researcher tested on his dataset unless he chooses to publish it. And there is a strong incentive to obscure this information. If a researcher releases the fact that he tested 20 hypotheses before finding 1 which was statistically significant, readers may discredit the result, or reviewers may reject it for publication. And if the researcher spent a lot of time and money collecting his dataset, it would feel like a waste to give up on the data just because his first five hypotheses tested on it don&#8217;t pan out. It might take a lot of restraint to not just keep testing hypothesis after hypothesis until he finds something publishable.</p>
<p>But even if researchers were excessively careful, that wouldn&#8217;t fully resolve the problem. When a hypothesis is confirmed by a dataset, we must consider whether it is truly a confirmation of the hypothesis being tested, or a result of the fact that 20 researchers tested 20 false hypotheses, and this one of the 20 happened to seem true by chance. That is, if enough hypotheses are tested over all, we may find a large number of false hypotheses among them that just happen to seem true.</p>
<p>What makes this problem more pernicious is that when a hypothesis fails to pan out, the result is often not published. This is due to the fact that hypothesis disconfirmations (e.g. &#8220;no association was found between cabbage eating and longevity&#8221;) are generally less interesting and harder to publish than confirmations (e.g. &#8220;an association was found between cabbage eating and longevity&#8221;). But since most new hypotheses in science turn out to be false, we should expect the number of negative results to be very large (except in situations where previously well validated results are being confirmed). Hence, the number of published test results will be much less than the number of total tests conducted, with test failures substantially underreported. So there is no good way to tell how many times a hypotheses failed to be confirmed by tests before one researcher finally ran one that seemed to confirm it. And if a very large number of false hypotheses are tested, but mostly just the ones that turn out to look true are published, you could end up with a field&#8217;s journals being flooded with false but seemingly verified hypotheses. In exploratory fields where almost all hypotheses are false, and where disconfirmations of a hypothesis are almost never published, you might even get into a situation where <a href="http://www.plosmedicine.org/article/info:doi/10.1371/journal.pmed.0020124">most published research findings are false</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2011/10/testing-too-many-hypotheses/feed/</wfw:commentRss>
			<slash:comments>6</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">249</post-id>	</item>
		<item>
		<title>Surprised? Update your model.</title>
		<link>https://www.spencergreenberg.com/2011/09/surprised-update-your-model/</link>
					<comments>https://www.spencergreenberg.com/2011/09/surprised-update-your-model/#comments</comments>
		
		<dc:creator><![CDATA[Spencer]]></dc:creator>
		<pubDate>Sat, 24 Sep 2011 18:01:26 +0000</pubDate>
				<category><![CDATA[Essays]]></category>
		<category><![CDATA[belief]]></category>
		<category><![CDATA[knowledge]]></category>
		<category><![CDATA[probability]]></category>
		<category><![CDATA[surprise]]></category>
		<category><![CDATA[truth]]></category>
		<category><![CDATA[update]]></category>
		<guid isPermaLink="false">http://www.spencergreenberg.com/?p=228</guid>

					<description><![CDATA[In order to make predictions, your brain must have a model of reality. This model is necessarily much simpler than reality itself. To see why, imagine that you are about to drop a baseball from waist height. Your brain can&#8217;t possibly know enough about the atoms composing that baseball and the air around it to [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In order to make predictions, your brain must have a model of reality. This model is necessarily much simpler than reality itself. To see why, imagine that you are about to drop a baseball from waist height. Your brain can&#8217;t possibly know enough about the atoms composing that baseball and the air around it to simulate what will happen at the atomic level. And even if your brain did have accurate knowledge about the atoms, using information at such a fine a level of detail would be extremely computationally inefficient. Instead, to predict the movement of the ball, your brain uses something closer to the model &#8220;when I drop an object, it will fall in a straight line down towards the ground, perhaps rotating slightly as it falls, and bouncing when it lands.&#8221; This is an accurate model that is far, far simpler than one that takes into account every atom.</p>
<p>Surprise is the emotion you feel when a prediction made by your brain&#8217;s model of what you are going to experience deviates substantially from what you actually experience. If you discover a large mark on your hand that you didn&#8217;t know was there, you will feel surprised. Your model of your hand did not match the visual experience produced by the light reflecting off of your hand. If you drop a baseball and it flies towards the ceiling, you will be even more surprised. Your model of what happens when you drop an object was just massively violated.</p>
<p>Our brain is making predictions constantly. Most of the time, these predictions are pretty accurate. When our sensory experience is in line with what our brain expects, it not only does not produce the feeling of surprise, but usually does not even push information about what was just observed into our conscious awareness. Making predictions and then checking these predictions against experience acts as a filtering mechanism. Since we can only keep a very limited amount of information in conscious awareness, it makes sense not to waste this space by filling it with boring stuff that your brain already predicted was going to happen. From the point of view of the survival of our genes, surprising experiences (i.e. ones that your brain mispredicted) are much more useful to reflect on consciously than unsurprising ones.</p>
<p>One thing to note about our brain&#8217;s predictions is that they do not correspond to a single, exact pattern of sensory experience. When you are in the forest, and you turn to the left, your brain cannot possibly predict the precise layout of leaves you will now see. But it does expect that you will see leaves, and that they will be (roughly speaking) within a certain range of colors, a certain range of sizes, and so forth. As long as this turns out to be true, you probably won&#8217;t even think about the fact that you are now looking at a different pattern of foliage than you were before. But if instead of trees there is a man standing next to you that you did not expect, you&#8217;ll immediately become consciously aware of that man.</p>
<p>By taking note of surprise when it occurs, we can learn to have truer beliefs about the world. To see how, imagine that you just heard about the <a href="http://en.wikipedia.org/wiki/Asch_conformity_experiments">Asch conformity experiments</a>, and you find yourself surprised to learn that most people in the experiment publicly gave the wrong answer to an obvious question just because a large number of other people publicly gave this wrong answer before them. Your surprise signifies an opportunity to know something that you didn&#8217;t before, and increase the accuracy of the beliefs in your model of reality.</p>
<p>But how should your model be updated based on this information? A naive updating would be to add a fact to your set of beliefs such as &#8220;If 35 people incorrectly claim that one line is bigger than another, then a large percentage of people will publicly claim the same belief as the group.&#8221; The problem is that when added in this way, the fact has not truly been incorporated into you web of beliefs. It has merely been glued to the web as a dangling outlier, without altering the web itself to naturally accommodate the new information.</p>
<p>To usefully incorporate surprising pieces of information, ask yourself:</p>
<ul>
<li><strong>Why did I find this information surprising?</strong> <span class="Apple-style-span" style="color: #808080;">This can help you hone in on what parts of your model need updating.</span></li>
<li><strong>What prior beliefs of mine does this information contradict?</strong> <span class="Apple-style-span" style="color: #808080;">Your credence in these prior beliefs should then be lowered.</span></li>
<li><strong>Which prior beliefs of mine does this information bolster</strong>? <span class="Apple-style-span" style="color: #808080;">Your credence in these prior beliefs should be increased.</span></li>
<li><strong>Which beliefs of mine are most natural to tweak so that this fact is no longer surprising?</strong> <span class="Apple-style-span" style="color: #808080;">Once the new information no longer seems surprising, it means it has been absorbed into your belief network.</span></li>
</ul>
<p>If you find the results of the Asch experiments surprising, it may be that you underestimate what people will be willing to do to avoid looking like an outsider. Or it may be that you underestimate the extent to which people will doubt their own sensory experience when it is contradicted by the opinions of others. In either case, the information you learned from the experiment should alter what you think about things besides what will happen in that precise experimental setup. For instance, this new knowledge may update your views about how cults create agreement among their members. Or it may make you realize that if dissent is not encouraged within an organization, there is a danger that conformity will take over, and the entire group may eventually succumb to obviously wrong, but socially reinforced beliefs.</p>
<p>If after you have updated your model the information you have learned still seems surprising, this indicates that your beliefs need further tweaking to accommodate the new information. You should not be surprised by the same (or very similar) information twice. If a friend cancels plans on you three times in a row, and you didn&#8217;t anticipate that they may cancel that third time, you failed to sufficiently process the first two occurrences.</p>
<p>The next time you find yourself feeling surprised, remember that it is a valuable opportunity to update your beliefs.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://www.spencergreenberg.com/2011/09/surprised-update-your-model/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">228</post-id>	</item>
	</channel>
</rss>
