Newcomb's Paradox

I agree, but it at least defeats the notion that outcome alone is sufficient. And this is important, because now you have updated your definition of what makes a strategy rational.

Great! I now introduce you to The Generalized Newcomb Game.

The Generalized Newcomb Game

  1. Participants do not know about this game, even in theory, before entering the room.

  2. Before the participants have entered the room, the predictor has predicted whether they will employ Strategy A or Strategy B.

  3. In the room, there is a Task T that can only be completed using Strategy A or Strategy B.

  4. If the predictor predicted that the player would use Strategy A, then it has put aside $1M for the player. This money is untouchable, it CANNOT be withdrawn after this. You can think of it as having already been sent to the player’s bank account, but the player has not been notified of this.

  5. The predictor truthfully explains all of the above, and the player knows for a fact that the predictor is 99% accurate. The player does not know if they have won the $1M however.

Now, Stategy A can literally be anything. For every choice of a task T, combined with each choice of a pair of allowed strategies A and B, we have specific premises creating specific games.

So, what if we chose Strategy A to be an irrational strategy, and Strategy B to be a rational strategy?

Such a choice is possible. Me and you can probably find a task T and a strategy A, such that we’d agree that strategy A is an IRRATIONAL way to complete the task T.

I mean, we both agree that irrational strategies exist, we just don’t always agree on which ones are which.

Now, the key is that the Generalized Newcomb Game does not make Strategy A rational in the case of playing that game. There is no pre-planning or iteration here. So, the irrational Stategy A is just irrational, even in that game. Simple as that.

Applying Strategy A to the Task T does not win the $1M! Those have already been won or lost, because the predictor has already made its prediction and set aside (or not) the money accordingly.

And yet, most players who use Strategy A will win $1M. It is a trivial fact that in the Generalized Newcomb Game, Strategy A is correlated with winning $1M. The Generalized Newcomb Game is a game with no deception from the predictor!

Therefore, we can apply your definition:

In my scenario, it is assumed that Strategy A is irrational. And yet, in the Generalized Newcomb Game, players using Strategy A will usually be millionaires, and players who don’t, won’t.

We thus derive this:

Strategy A is irrational AND Strategy A is rational.

That is a contradiction. Therefore, your definition of rationality is not usable. We need a subtler definition, because even in a game that doesn’t lie to you, you still CANNOT look at outcome alone.

You have to see how the outcome relates to controllable and uncontrollable factors. This has to be your definition of rationality, because it avoids contradictions and allows us to make usable, rational decision theories.

Here’s my definition:

Strategy A is rational for task T if APPLYING it improves your outcome more than all other strategies, assuming it is being employed in a non-deceptive game.

The key word here is applying it. If Strategy A is non-causatively correlated with better outcomes that other strategies because the PAST, uncontrollable predisposition to use Strategy A was rewarded by the game’s construction… then that is not going to be a strategy labelled as “rational” by my definition above.

You can’t just assume this as it begs the question. You need to prove that some strategy is both irrational and most profitable.

But I’m confused, because your two-boxer reasoning uses the same justification; it is most rational because “it wins me $1,000 more”.

So we appear to agree on what it means for a strategy to be most rational. I am simply showing that your application of it is myopic and misleading. You need to compare the winnings of one-boxers and two-boxers, and you’ll see that one-boxing almost always wins you $999,000 more.

I think you’re reading “collapse two nodes” as a claim about temporal identity as though I’m saying the predisposition and the choice happen at the same time. I’m not. They obviously occur at different times. What I’m saying is that they don’t constitute two separate causal handles: two independent levers that the agent (or the world) can operate on independently. (I did acknowledge “It marks, at most, a temporal cut […]”)

Consider an analogy: knowing French and speaking a French sentence occur at different times, but a diagram that placed them as separate nodes linked by a probabilistic causal arrow (“knowing French ontically determines, at 95%, that the speaker will produce a grammatically correct French sentence”) would fundamentally misrepresent the relationship. The knowledge isn’t a prior event that causes the speech as a downstream effect. It is the standing capacity whose exercise is the speaking. You can temporally distinguish the acquisition from the deployment without treating them as two independent causal relata.

Similarly, the agent’s predisposition and the agent’s choice are temporally distinct but not causally independent in the way your diagram requires. The predisposition is not an event in the past that pushes the choice into existence at 99% reliability. It is a rational capacity that the agent exercises in the act of choosing. Your diagram treats them as two handles, which is what generates the Phase 2 reasoning: “The predisposition handle has already been pulled (Phase 1); now I must optimize with the choice handle (Phase 2).” But there is only one handle: the agent’s rational agency exercised over time.

As for P-determines: I haven’t given it up. The node-collapsing point and the R(g, e₁, e₂) formulation are complementary. The rational ground (g), which is what your diagram splits into “predisposition” and “choice”, co-determines two genuinely separate events: the predictor’s anticipatory action (e₁) and the agent’s action (e₂). Those two events are separate nodes, at different times, with no causal arrow running between them. What connects them is their common dependence on the agent’s rational grounds.

And the 1% failure rate comes from the predictor’s imperfect sensitivity to those grounds, as your own diagram already shows. On your picture, the fallibility is located at the predictor’s prediction node, not at the arrow between predisposition and choice. Collapsing those two nodes into a single rational capacity changes nothing about where the fallibility lives. The predictor is still a separate agent making an imperfect anticipatory judgment about what the agent’s rational grounds will yield and that is where the 1% sits on both our pictures.

Nozick doesn’t explicitly argue for one-boxing in The Nature of Rationality. What he does is something more interesting. He abandons the confident two-boxing of the original paper and develops a “decision-value” framework that weights causal and evidential considerations, without arriving at a stable resolution. He even considers the possibility that CEU (causal expected utility) could recommend that a rational agent swallow a pill that would cause them to employ EEU (evidential expected utility) literally: a magical pill that rewires their decision-making. This is strikingly similar to @AlveK’s prescription that agents should find ways to make themselves one-box in the iterated version of the problem. In both cases, the causal principle concedes that its own application requires a prior intervention that undermines its authority as a universal principle of rationality. If CDT can recommend making yourself into someone who doesn’t follow CDT, something has gone wrong with CDT’s claim to be the correct theory of rational choice.

On your three claims: I accept claims 1 and 2. Once the prediction is fixed, two-boxing yields $1,000 more. That’s just dominance reasoning conditional on a fixed prediction, and it’s true. The issue is claim 3. “The prediction has already been made and I can choose, therefore me taking two boxes will win me more money.” The “therefore” assumes that the prediction being temporally prior makes it causally independent of the choice. But that is the very thing at issue between us. As I’ve argued, the prediction’s being settled in time does not make it settled independently of the agent’s rational grounds. The predictor made their prediction because they anticipated how the agent would deliberate. So the move from “the prediction has already been made” to “therefore my choice is causally inert with respect to it” is not a valid inference. It is the CDT assumption that I’ve been challenging throughout this thread.

If CDT can recommend making yourself into someone who doesn’t follow CDT, something has gone wrong with CDT’s claim to be the correct theory of rational choice.

Not really. Here’s a game:
“If you follow CDT, you lose $10, otherwise you win $10”

CDT would say to not follow CDT but that’s because the problem demands it. This could be constructed for any theory of rational choice.

The “therefore” assumes that the prediction being temporally prior makes it causally independent of the choice.

Right, so you are saying your choice influences (causally) the prediction?

Your $10 game has the structure of the liar’s paradox. The payoff depends on whether the agent follows CDT, but CDT’s recommendation depends on the payoff. There is no stable assignment of outcomes. No game master could actually fulfill the stated obligation. This is quite different from the Newcomb problem, which is perfectly well-specified: the predictor predicts, fills the box, the agent chooses. There is no self-referential loop, only a disagreement about the correct strategy.

On causation: no, I am not saying my choice causally influences the prediction. That would be retro-causation, which I’ve explicitly rejected throughout. What I’m saying is that I lack the power to make my choice and the prediction causally independent, because they share a common source in my rational grounds. The predictor anticipated my reasoning; my reasoning produces my choice. I cannot choose in a way that severs this connection, not because I lack freedom, but because any rational basis for my choice is precisely what the predictor was sensitive to.

What the conditions of the experiment do establish is a perfectly normal forward-directed causal connection between my choosing to one-box and my walking out (most of the time) with $1,000,000 mediated by the predictor’s anticipatory sensitivity to my rational grounds. The causal path runs from my rational grounds, through the predictor’s anticipation, to the filling of the box, and separately from my rational grounds through my deliberation to my choice. Both paths run forward in time. The result is that one-boxing and walking out with $1,000,000 are causally connected, not because my choice reaches back to change the box, but because the predictor has already ensured that the box contents answer to the rational grounds on which I act.

Your $10 game has the structure of the liar’s paradox. The payoff depends on whether the agent follows CDT, but CDT’s recommendation depends on the payoff. There is no stable assignment of outcomes. No game master could actually fulfill the stated obligation

What do you mean? Someone following EDT would win the $10 dollars. CDT would say to not follow CDT and you would win the $10 dollars. There is no liar’s paradox. It’s a perfectly well-defined meaningful game.

This is quite different from the Newcomb problem, which is perfectly well-specified: the predictor predicts, fills the box, the agent chooses. There is no self-referential loop, only a disagreement about the correct strategy.

I never compare the game to Newcomb problem. I just wanted to show that your claim that

is nonsense.

Outside of that, your words are not making clear at all what you believe in or why you think one should OB. The only sense in which I could understand what you are saying is through a perfect predictor or a lack of choice. If I say you do assume that, you deny and keep talking as if it were the case.

I like them separate. What I don’t like is the epistemic arrows not having any tail. There no ‘X causes you to know Y’ since there’s no X anywhere, just all Y’s.

So for instance, player’s choice epistemically determines (99%) the prediction. Strangely enough, the deliberation (not shown as a separate circle) does not in any way do this. The opening of the box epistemically determines (100%) box 2’s contents. That should be a separate arrow, but its all one (two actually) arrows, all heads and no tails.

What if my predisposition is to mess with the predictor? What does it predict then? Such a predisposition probably doesn’t care about the money.
Speaking of that, in the OP, you have a vote for iterated version. Assuming the ‘fixed amount of time’ is short, then one-boxing (the one with just 1000 in it) is usually the rational choice. Mathematically, unlimited thousands is no more than unlimited millions, so opening a potentially empty box is more work than it’s worth when all I wanted was to pay my meal tab. There was no vote option to select this one.

Clarification: What Causal and Evidential Decision Theory Ought to Recommend

Since the discussion has branched in several directions, I want to state my position as clearly as I can, because I think it has been misread by some participants as a defence of evidential decision theory. It is not.

I am a one-boxer. But my grounds for one-boxing are causal, not evidential. Here is the core of the disagreement as I see it.

CDT says: choose the action that causes the best outcome. I accept this principle. What I reject is the assumption, standard in the decision theory literature, that the only causation available to the agent in the Newcomb setup is efficient event-event causation operating forward from the moment of choice. Under that assumption, the agent’s deliberation is causally inert with respect to the already-settled box contents, and two-boxing follows straightforwardly.

EDT says: choose the action that is the best evidence of a good outcome. This recommends one-boxing, but for the wrong reason. It treats the connection between the agent’s choice and the box contents as merely evidential: one-boxing is good news about what’s in the box. On my view, the connection is genuinely causal, and misdiagnosing it as evidential has a serious cost: EDT cannot properly distinguish the standard Newcomb case from medical Newcomb cases (such as the smoking gene scenario), where two-boxing is correct.

My view: The predictor in the standard Newcomb case is reliable because they are sensitive to the agent’s rational grounds: the very grounds on which the agent will act. The predictor serves, in effect, as a causal facilitator: they extend the reach of the agent’s practical determination into a domain (the prior filling of the box) that the agent’s own hands cannot reach but that their rational grounds can, via the predictor’s anticipatory sensitivity to them. The agent’s deliberation is therefore genuinely causally efficacious with respect to the box contents, not through retrocausation but through the ordinary forward-directed structure of rational agency as anticipated by a perceptive predictor.

This is rational causation: the agent, in determining what to do on rational grounds, brings about the outcome (including the part of the outcome the predictor has already arranged in anticipation of those very grounds). CDT, properly understood with an enriched conception of what agents can cause, recommends one-boxing.

For distinguishing cases where one-boxing is correct from cases where two-boxing is correct is simply this: what is the predictor (or the correlation) tracking? If it tracks the agent’s rational grounds — the exercise of their deliberative agency — then the agent’s deliberation is causally efficacious, and one-boxing is correct. If it tracks a brute disposition that operates independently of the agent’s reasoning (as in the smoking gene case, where the gene causes both the propensity to smoke and the cancer), then the agent’s deliberation is causally inert with respect to the correlated outcome, and two-boxing is correct.

Both CDT and EDT, as standardly formulated, miss this distinction. CDT misses it because it recognizes only efficient event-event causation, EDT because it treats all evidential relevance alike. The resources for drawing the distinction come from the philosophy of action: specifically, from the recognition that rational agency is a genuine causal power, not reducible to event-event efficient causation, and that practical knowledge, in Anscombe’s phrase, is “the cause of what it understands” (Intention, §48).

1 Like

Absolutely not. All I assumed is that I could find a strategy we both agree is irrational. I can find one.

But then if I set Strategy A as this irrational strategy, and I set Strategy B as some rational strategy, and I input them into the Generalized Newcomb Game, then I get a situation where Strategy A, despite being irrational, wins more money than the rational Strategy B.

Under your definition, this makes Strategy A rational. But we already assumed it was irrational, based on the fact that I could find any irrational strategy for Task T. And this trick works with any irrational strategy, so it does not need to be specific.

But I get that this is easier to understand if we look at an example.

So, let’s choose a Task T and a Strategy A and a Strategy B. I will only proceed with my argument once you agree that Strategy A is irrational to use for solving Task T, and that the Strategy B is rational. If we have to go through a few possible choices, then so be it. Please feel free to propose examples.

A Possible Example:

The Task T:

You enter the room and there is a table. You are given a pack of 52 cards. Your task is to place all the cards on the table (the cards being on the table means the table supports them). For every card that directly touches the table, it costs you $1.

In this game, there are only two strategies allowed to solve this problem.

Strategy A:

You take out every card and place them on the table individually, never letting a card be on top of another card.

Strategy B:

You simply place the whole pack of cards, with the packaging intact, on the table.


These are, by the rules of the game, the only allowed strategies.

Strategy A is obviously the irrational choice for task T, because it will cost you $52. Strategy B is obviously the rational choice for task T, because it will cost you $0.

If you agree on this, then we can move on. If not, let’s find a new example. However, a cost of $0 is better than a cost of $52, so I’m not sure how you could disagree.

Inputting the Examples into the Generalized Newcomb Game

Let us instantiate the variables of Task T, Strategy A and Strategy B. We then get the example game here:

Example Game 1:

  1. Participants do not know about this game, even in theory, before entering the room.

  2. Before the participants have entered the room, the predictor has predicted whether they will employ Strategy A or Strategy B.

  3. In the room, there is a Task T that can only be completed using Strategy A or Strategy B.

  4. Task T is to put all 52 cards on the table. For every card that directly touches the table, it costs the player $1.

  5. Strategy A is putting all 52 cards on the table individually, making them all directly touch it.

  6. Strategy B is putting the pack of cards on the table, keeping the packaging on.

  7. If the predictor predicted that the player would use Strategy A, then it has put aside $1M for the player. This money is untouchable, it CANNOT be withdrawn after this. You can think of it as having already been sent to the player’s bank account, but the player has not been notified of this.

  8. The predictor truthfully explains all of the above, and the player knows for a fact that the predictor is 99% accurate. The player does not know if they have won the $1M however.

Strategy A is clearly irrational. And yet since the predictor is quite accurate, Strategy A will be correlated with winning $1M - $52.

We have already determined that Strategy A is irrational. Now, we see that since Example Game 1 rewards it, it may be… rational after all? Well, let’s look at your definition:

In this case, you know all the rules. And also, being someone with Strategy A will almost always win you more money.

Therefore, Strategy A is rational, per your definition.

But we also started out with Strategy A being irrational. That is a contradiction.

How My Definition Deals With This

Strategy A is rational for task T if APPLYING it improves your outcome more than all other strategies, assuming it is being employed in a truthful game.

Well, in Example Game 1, applying Strategy A means I incur a cost of $52, as opposed to $0.

It is not applying Strategy A that wins me the $1M. The predictor gave me those $1M before the decision-making even begun. It gave me those, or did not, based on its past, now unchangeable prediction that I would, or would not, employ Strategy A.

My definition easily deals with this example game, because I differentiate between two things:

Is Strategy A non-causatively correlated with winning more, or is it causatively correlated with winning more? Only the latter can make it rational!

Applying Strategy A is causatively correlated with winning less, because it causes you to incur an unnecessary cost of $52. Therefore, Strategy A is irrational.

And yet it is causatively correlated with winning more in Example Game 1, because Example Game 1 is constructed to reward the players with the predisposition to use that strategy, and this rewarding happened before the decision-making even started.

Your definition is not capable of dealing with these example games. Can we agree?

If so, then consider this:

What if the Newcomb Game really is an example of the Generalized Newcomb Game… and what if OBing could be irrational, despite being correlated with higher winnings…

WHAT IF OB-ING IS ONLY NON-CAUSATIVELY CORRELATED WITH WINNING THE $1M?

This is the heart of Causal Decision Theory. We make decisions based on what we can cause. Applying our CDT strategy will always win us more than not applying it, but whether the CDT strategy is correlated with winning more depends on if the particular game happens to be constructed such that before we can even make a decision, we have been punished for having the predisposition to probably apply CDT.

You have to see the bigger picture here. We can construct infinitely many games to gives us an infinite diversity of results. The Newcomb Game is not actually special in its idealized construction, but it happens to be especially tricky and counter-intuitive.

We must look at the invariants. The rational decision theory is the one whose application always wins us more. Whether this means it is correlated with winning more depends on arbitrary details of the game’s construction that can be tailored to punish or reward anything before the player’s decisions even get made!

I am going to bed right now, but before I leave the discussion for now, I will ask you a question regarding what you seem to be saying here.

You seem to be saying the predictor is not merely sensitive to the player’s PRE-disposition at time t_1, but rather the player’s temporally extended disposition for XBing, which is temporally extended throughout the game?

Is that your view?

Roughly, though I should clarify a bit.

I’m not saying the predictor is sensitive to a disposition that is somehow “temporally extended” in the sense of being spread across both Phase 1 and Phase 2 as a process. The disposition is a standing rational capacity. It is fully present at the time the predictor makes their prediction, and it is stable through time: the same capacity that is present at the time of prediction is the one the agent will exercise when the moment of decision arrives.

Now, how the predictor tracks this is a separate question from what it is that they track. A Laplacean predictor could fulfill their role entirely by reading micro-physical states, recognizing that physical state P1 reliably leads to physical state P2, and that P2 is a realization of one-boxing or two-boxing. Such a predictor need not understand why the agent one-boxes; they need only anticipate that they will. The predictor can be entirely blind to the person-level significance of the states they read.

But from the agent’s own practical standpoint, the picture looks quite different. What explains the agent’s choice to one-box (i.e. what makes it intelligible as an exercise of rational agency rather than a mere physical outcome) is that their antecedent circumstances and dispositions (which happen to be realized by P1) make their choice (which happens to be realized by P2) rationally intelligible in light of them. The agent takes themselves to be actualizing a standing rational capacity in specific circumstances, and they recognize that the predictor, however they operate, has tracked the physical realization of precisely those rational grounds.

This is why the agent’s disposition need not be rigid or their character highly predictable. Whatever deliberative trajectory the agent actually follows — however tentative, conflicted, or capricious it may be — that trajectory is physically realized, and the predictor reads the physics. What matters for the rationality of one-boxing is not that the agent have an iron character, but that they understand their own choice as an exercise of practical reason whose physical realization is what the predictor was sensitive to. It is this self-understanding that gives the agent grounds for confidence that one-boxing will secure the large prize. It’s not a guarantee, since the predictor is fallible, but a rational expectation grounded in the recognition that their deliberation and the predictor’s anticipation are connected through the physical realization of common rational grounds.

Here is a schema meant to illustrate my previous post that I’m no longer able to edit:

So the physical state P_1 leads to an already predetermined physical state P_2 i.e. the player lacks agency. When you say to pick OB, what you are actually saying is that you want the physical state P_1 to reflect an agent who will OB.

The difference is that when I say pick OB/TB I am not trying to change the physical state P_1 because you can’t change it once you are in the room. In your view, once you are in the room, the agent simply follows what their past physical state leads them to do. So no choice/agency/freedom however you want to call it, there is no decision once you are in the room.

You need to understand that nobody even two-boxers is saying that prior to the prediction you should appear as an agent who TB. So if you interpret “what should you do in front of boxes” as “what should be your physical state before the prediction” then yes, you should OB even CDT says you should, but this isn’t the question two-boxers are answering.

We consider that the physical state in the past doesn’t predetermine the choice in the room otherwise obviously there is no choice in the room and you are forced to interpret the question as a completely different question.

No we haven’t.

We can make this even simpler. There are two lotteries, both with a $1,000,000 reward and a $1,000 consolation prize. I may only play one of the lotteries, and I may only buy one ticket.

  1. A ticket for lottery A costs $52, with a 99% chance of winning.
  2. A ticket for lottery B costs $1, with a 1% chance of winning.

You are arguing that playing lottery A is irrational because the ticket is more expensive. I am arguing that playing lottery A is rational because the expected return is greater.

Any rational person will play lottery A.

Or to make the above like the Newcomb problem:

  1. I must choose to buy a ticket for either $1 or $52
  2. If I am predicted to buy it for $1 then I win $1,000
  3. If I am predicted to buy it for $52 then I win $1,000,000
  4. The prediction is 99% accurate

Any rational person will buy the ticket for $52, because almost everyone who does wins $1,000,000 and almost everyone who doesn’t wins $1,000. This fact matters more than the myopic “the prediction has already been made and my reward set in stone, therefore I ought by the cheaper ticket” reasoning.

This is largely why I’ve stopped directly responding most of the time. We cannot move on because B is not the rational choice, as pointed out in several places, including the excellently worded recent posts by Pierre.

Your reasoning above only looks at the $52 and fails to take the larger utility into account. That’s irrational. The large prize is based neither on your choice (which would be backwards message sending, impossible under our physics), nor on the prediction, but rather on that on which the prediction was made, which is your predisposition. You cannot be predisposed to OB and then choose TB, so rationally, the only way to be an OBer is to OB. OB takes place at the time of choice, but being an OBer is a state that exists before the prediction is made, and it isn’t something that changes between prediction and choosing.

Any rational person will buy the ticket for $52, because almost everyone who does wins $1,000,000 and almost everyone who doesn’t wins $1,000. This fact matters more than the myopic “the prediction has already been made and my reward set in stone, therefore I ought by the cheaper ticket” reasoning.

"Any rational person avoids IQ tests, because almost everyone who avoids them doesn’t have psychological troubles and almost everyone who takes them does have psychological troubles. This fact matters more than the myopic “you already either have or don’t have psychological troubles, therefore taking or not taking the IQ test doesn’t change anything” reasoning.

You cannot be predisposed to OB and then choose TB,

Why do you (and Pierre too) add your own assumptions to the problem, change the fundamental question we try to answer and then act as if you were answering the original problem?

At least thank you for stating so clearly what you believe in. I am not against trying to solve different versions of the problem but why act as if you were answering the question “what to choose in front of the boxes?”? You clearly don’t believe there is any choice to have at this point, everything has been determined by the predisposition.

This comparison you keep trying to make with IQ tests and mental health is a straw man, as explained before.

Oh yeah, I forgot you just didn’t understand the problem and kept comparing it to scenarios that had nothing to do with Newcomb problem.