Newcomb's Paradox

So had I, circa 2008, before I read and thought more about the problem. What makes two-boxing appear compelling as the rational strategy is the dominance principle as it does apply to the smoking lesion problem, where evidential causal theory flounders, as Yudkowsky and Soares argue in the paper your referenced earlier. Hence, a satisfactory defense of one-boxing can’t be content with just pointing out how irrational the two-boxing strategy intuitively seems to be on account of the bad results they get. It must also explain why it is that the standard Newcomb problem can’t be construed as a variation of the smoking lesion problem where rational players do get worse results.

You start from before the prediction with your robot. If you had to program the robot after the money was already in the bank account (or not) then the most rational choice is to tell the robot to take the $1,000.

Why is this so hard to understand for you? Can you really not see the difference between having some choice before the prediction and after the prediction?

The reason why you think the smoking lesion problem is different is because you can’t equivocate anymore between before and after. You can’t interpret “should you smoke or not” as “should I have the gene or not” because it’s too different of a question. That’s it.

The smoking lesion (just like the IQ test problem I laid out) is entirely equivalent to the Newcomb problem and the only reason you say to OB in the Newcomb is because you keep changing the question into a different one (like Yudkowski) but you either don’t realize this or you do and you think that’s how everyone is supposed to see the question. I repeat, the question is about choosing after the prediction happened, i.e. with the past of the agent fixed. No, this does not mean the agent will do whatever they are predetermined to do, there is no reason to suppose determinism.

The worst part is misunderstanding the TB position (like Yudkowski). In your robot scenario, everyone and CDT will tell you to program the robot as to not take the $1,000. That’s not where the disagreement is.

The equivalent of your robot scenario in the smoking lesion is: “Imagine you can construct your robot and give it the gene or not. Well then it’s obvious that to avoid cancer, you should not give it the gene and therefore make it not smoke, so the answer is you should not smoke in the smoking lesion problem”. See how this is simply not answering the question?

You say that everyone agrees the programmer should program the android to walk away, and that the disagreement only arises after the prediction, when the agent faces the $1,000 with the bank transfer already settled. You say my thought experiment misses the real dispute because the programming happens before the prediction.

I don’t think it misses the dispute. I think it locates it with a precision that the standard formulation obscures. You’re right that there’s no controversy about what the programmer should do. CDT, EDT, everyone seemingly agrees (including @AlveK in the OP) that the android should be programmed to walk away. But the interesting question isn’t what the programmer should do. The interesting question is what the android itself should do when it walks into the room, faces the $1,000, and, on your account, finds itself in a world where the bank transfer is already settled. That’s indeed where the real disagreement lives.

Predictor versus retro-dictor from the android’s perspective.

Suppose we replace the predictor with a retro-dictor. Instead of examining the android’s decision procedure before it acts and filling the box accordingly, the retro-dictor examines the decision procedure after the android has acted, perhaps by merely inspecting the android’s hardware and code, and then fills or empties the box accordingly.

From the programmer’s standpoint, nothing changes between these two setups. Either way, you program the android to walk away. No controversy there. But from the android’s own temporally situated perspective, on your account, the two setups should differ crucially. With the predictor, the box contents are settled before the android acts. So you’d say the android’s walking away is causally inert with respect to the box. With the retro-dictor, the box contents are not yet settled when the android acts. The retro-dictor will fill the box after examining the android’s dispositions. So even on your view, in the retro-dictor case, the android’s walking away is straightforwardly causally efficacious. The box-filling hasn’t happened yet; it will be determined by what the retro-dictor reads in the android’s decision procedure.

But here’s the thing: from the android’s own practical standpoint, the two setups are indistinguishable. In both cases, the android’s dispositions are within the purview of the box-filler. Whether those dispositions are read in advance by a predictor or examined after the fact by a retro-dictor, the box contents will reflect them. The android programmed to walk away will find $1M in the box regardless of whether the box was filled yesterday or will be filled tomorrow. The android programmed to grab the $1,000 will find the box empty either way. For the android, standing in the room, deliberating about what to do, the temporal ordering of the box-filling is practically idle. What matters is that its dispositions are transparent to the box-filler and this is equally true in both cases.

So the temporal distinction you’re relying on between “before the prediction” and “after the prediction” does no practical work from the agent’s own perspective. This distinction collapses the moment you see that what determines the box contents is not when the box-filler acts but what they’re rationally sensitive to when they deliberate what to do. Those are the agent’s rational dispositions, whenever those dispositions have been already read by the predictor of not.

The only way to maintain that the predictor case and the retro-dictor case differ from the android’s own standpoint is to imagine that the android could somehow act in a way that is decoupled from its own programming. In the retro-dictor case, the android can’t get away with this, because the retro-dictor will examine the actual decision procedure after the fact and catch any deviation. But in the predictor case, you might think, the prediction is “already locked in,” so the android could grab the $1,000 and the predictor wouldn’t know.

But what would this mean? It would mean the android acting contrary to its own programming and executing a decision that its decision procedure doesn’t produce. This is incoherent for an android since the programming is what generates the android’s action. There’s no gap between the program and its execution where a different choice could slip in. The android that grabs the $1,000 is an android whose program says to grab the $1,000 and the predictor, modelling that program, will have anticipated it.

What you’re imagining, in effect, is a kind of bank heist: the agent exploiting a gap between their modelled dispositions and their actual behaviour, acting outside the predictor’s purview by acting outside the purview of their own rational character. You’re imagining a immaterial ghost reaching in from outside the world to pull the android’s hand in a direction independent of its programming.

Regarding the special case of the smoking lesion

The reason I say one-box in Newcomb and “go ahead and smoke” in the smoking lesion is not that I’m equivocating about timing. It’s that I’m asking: what does the correlation track?

In Newcomb, the predictor tracks the agent’s rational dispositions that drive the agent’s deliberation about what to do. And this is why the two-boxer’s question — “what should I rationally do, given that my rational dispositions are what they are?” — is incoherent. You cannot deliberate about what to do in a way that isn’t an exercise of those very dispositions. Any answer you arrive at is their actualization. There is no standpoint outside your own rational character, needs and world knowledge, from which to assess what you “really” ought to do, because your rational character is what does the assessing.

In the smoking lesion, the gene tracks a brute biological disposition: an arational urge to smoke that operates independently of the agent’s rational assessment of what it is that, all things considered, they ought to do. And this is precisely why it does make sense to ask: “Given that I possibly have this irrational disposition to smoke, what should I rationally do?” The agent’s deliberation is not an exercise of the gene. The gene pushes from behind; reason assesses from above. You can coherently step back from the disposition and evaluate whether to act on it, because the disposition and the rational assessment are different things operating through different channels.

That’s the real asymmetry. In Newcomb, the disposition tracked is the agent’s rational agency, so you can’t deliberate your way outside it. In the smoking lesion, the disposition tracked is arational, so you can (and should) deliberate independently of it.

Are you still assuming your determinism?

Assume no determinism and the agent has agency in front of the boxes despite a fixed past, should the agent pick one box or two boxes?

I’m not assuming determinism. In the case where determinism is false, and the agent lacks a firm rational predisposition to one-box or two-box, the predictor’s accuracy could drop considerably. But even a predictor with limited accuracy yields a higher EV for one-boxing than for two-boxing, so long as the residual correlation stems from the predictor’s ability to track the agent’s rational dispositions, which is something that some two-boxers like yourself readily grant but construe as a setup that rewards irrationality.

What matters is what the correlation tracks, not how strong it is. The one-boxing recommendation follows from the predictor’s sensitivity to rational dispositions, not from determinism.

You also ask me to assume the agent “has agency in front of the boxes despite a fixed past.” That “despite” is doing a lot of work. It seemingly assumes that the fixed past is something the agent’s agency must overcome and that genuine agency requires somehow acting independently of one’s prior rational formation. But on my account, the agent’s rational dispositions are part of what the fixed past consists in, and exercising them is the agency, not an obstacle to it. Their own past upbringing, character formation, memories, etc., aren’t things the agent acts independently from. Rather, they constitute the rational character through which they act.

So, no, I’m not relying on determinism. I am only assuming that it’s not possible to rationally assess one’s own practical situation (i.e. what it is that one ought to do as an embodied rational agent) entirely from outside one’s own rational character and dispositions. Unless, perhaps, you hold the rational soul to be something immaterial that is forever outside the purview of the predictor and unconstrained by the agent’s own embodied history and cognitive abilities. But that amounts to denying the possibility of the standard Newcomb setup being implemented in our world.

In the case where determinism is false, and the agent lacks a firm rational predisposition to one-box or two-box, the predictor’s accuracy could drop considerably.

Again, rejecting determinism doesn’t imply the agent lacks a firm rational predisposition to anything nor does it imply that the predictor can’t be accurate.

That “despite” is doing a lot of work. It seemingly assumes that the fixed past is something the agent’s agency must overcome

No it doesn’t.

that genuine agency requires somehow acting independently of one’s prior rational formation

Again, this was never stated by me anywhere

Their own past upbringing, character formation, memories, etc., aren’t things the agent acts independently from. Rather, they constitute the rational character through which they act.

I don’t care. Assume they have agency in front of the boxes, i.e. reject determinism, that’s all I am asking.

I am only assuming that it’s not possible to rationally assess one’s own practical situation (i.e. what it is that one ought to do as an embodied rational agent) entirely from outside one’s own rational character and dispositions.

Right. So I am asking again to be clear, you think one should OB then?

Do you agree with the fact that for any agent in front of the boxes, TB gives them $1,000 more?

Indeed! You make the important concession that rejecting determinism doesn’t imply the agent lacks firm rational dispositions, nor does it imply the predictor can’t be accurate. But notice that this gives my account everything it needs. My one-boxing recommendation rests on the predictor’s sensitivity to the agent’s rational dispositions. I don’t need determinism. I just need stable enough dispositions and a predictor who can track them with some degree of reliabillity.

But then you ask me to “assume the agent has agency in front of the boxes, i.e. reject determinism.” Look at what that “i.e.” is doing. It equates having agency with rejecting determinism, as though they were the same thing. But you just conceded that rejecting determinism leaves the predictor’s accuracy and the agent’s stable dispositions intact. So what is “agency” adding here, beyond what’s already captured by “the agent exercises their stable rational dispositions”? The only distinctive content it could have is the ability to act independently of those dispositions, which is exactly the bank heist picture from my previous post. This is the agent’s soul pulling the agent’s hand from outside the material world.

You say that you never claimed genuine agency requires acting independently of one’s prior rational formation. But when I offer a positive account of what agency does consist in (i.e. the exercise of one’s rational character as constituted by one’s past formation, dispositions, and cognitive capacities) you respond with “I don’t care.” If you don’t care what agency consists in, then how can the demand to “assume agency” do any work in the argument? If agency is the exercise of one’s rational character, one-boxing follows. If agency requires something beyond the exercise of one’s rational character and cognitive abilities, I need you to tell me what that something is.

Finally, you ask whether I agree that “for any agent in front of the boxes, two-boxing gives them $1,000 more.” This is presented as a bare accounting fact. But it presupposes the CDT picture rather than establishing it, and it does so in a way that reveals an internal tension. For the dominance argument to work, you need the agent to be powerless to determine the contents of the opaque box (that’s already settled by their prior dispositions as the predictor tracked them). But you simultaneously need the agent to be powerful enough to secure an extra $1,000 by deciding to two-box, which means deciding to be something other than what their dispositions (as tracked by the predictor) already determined them to be.

The agent is powerless to influence the box contents through their rational character, yet powerful enough to override that same rational character by reaching for the extra $1,000. These two claims are consistent only if the power to two-box comes from the unpredictable ghost that pulls the hand after the predictor has done their job rather than from the agent’s rational dispositions. On my view, by contrast, the agent has no need for an ability to act outside of the purview of the predictor. They’re perfectly happy for the predictor to be able to look over their shoulder. The predictor being able to foresee what they will do (one-box or two-box) takes nothing from their ability to decide on rational grounds how they wish the opaque box to be filled. On the contrary, it enables them to secure the large reward.

There is a kind of “spookiness” to the predictor that might be affecting people’s judgements. I wonder how people would respond differently if just the character of the predictor was altered a bit.

Suppose the predictor were an AI that integrates multiple factors: demographics, social media posts, physical features. It uses these to produce a prediction which is not absurdly accurate, but still reasonable: say 65% accuracy.

Given this setup, would OBs still choose to be OBs? The improved odds are still considerable, so that the expected outcome is much higher to be OB. Yet, the predictor has no special, “spooky” insight into your decision making itself. I find OB much harder to justify in this scenario.

I have the opposite intuition. I find two-boxing much easier to justify in this scenario, and I think the reason why helps us see where the one-boxing recommendation actually comes from.

Your AI predictor integrates demographics, social media posts, and physical features to produce a 65% accurate prediction. But notice what it’s tracking: raw behavioural tendencies that are the kind of brute dispositions that operate independently of the agent’s present rational deliberation. The predictor modelled your social media history, not the reasoning you’re doing right now in front of the boxes. This makes the scenario structurally closer to the smoking lesion (or to Suny’s IQ test variation) than to the standard Newcomb case. The correlation between your choice and the prediction runs through a common cause (your demographic/dispositional profile) that the correctness of your present deliberation is genuinely independent of.

And this is why the EV calculation, despite being arithmetically correct at 65%, is misleading. In the standard Newcomb case, you can’t rationally two-box without this constituting a rational disposition to two-box that the predictor tracked. Your deliberation and the predictor’s model are sensitive to the same rational grounds. But with the demographic AI, you can rationally two-box without undermining the basis of the prediction, because the prediction wasn’t based on your rational grounds in the first place. It was based on the statistical tendencies of people who share your profile. Your present reasoning is invisible to it. So the agent can coherently reason: “People like me tend to one-box 65% of the time, but I can see that two-boxing is the right choice in this situation” and this reasoning doesn’t undercut the prediction, because the prediction never modelled this reasoning.

The point of practical reason is to determine what you rationally ought to do, not to discover what you’re statistically likely to do. It would be irrational to reason: “If I two-box then I probably have the tendency to two-box, so I should one-box as evidence that I don’t.” That’s treating your own deliberation as mere evidence about your dispositional profile, which is exactly the EDT error on my view. What makes one-boxing rational in the standard Newcomb case is not that it provides evidence of a favourable prediction, but that the predictor is sensitive to the very rational grounds on which you’re making the choice in the present situation. Remove that sensitivity and the one-boxing recommendation loses its justification.

Argh. I agree. I’m an idiot, I swapped my acronyms in my post.

Ah! Got it. No worries. Your mistake triggered me into thinking about some interesting complications that arise from the method used by the predictor to assess whether the agent will one-box or two-box. There may be interesting cases that are somewhat intermediate between the medical and Laplacean variations of Newcomb’s problem. So, it’s all good.

Agreement aside, I still think you are fooling yourself (though I’m late to this interesting thread, and still working it through). As of now I think @AlveK got the whole thing right in his op, and that the only way out is something like a legal contract not to TB, in advance. Or, to fool yourself into having your cake and eat it too: to convince yourself that it is rational, when actually confronted by both boxes, to choose only one. The essence of the dilemma is, to me, that this is strictly irrational.

Here is a question for you: suppose there are no boxes and the money is in plain sight. Would you still “OB”? If not, why do the boxes matter to your decision?

For the predictor to fulfil their function of filling the opaque box with $1,000,000 if and only if the agent is poised to one-box, it is necessary that the box remains opaque until the agent makes their decision. Otherwise, you introduce a feedback loop: the predictor must factor in the consequences of their own prediction into the data on the basis of which they make it (since the agent sees what was predicted before choosing), and that’s impossible. So the predictor simply can’t do their job with transparent boxes. If the predictor is disabled, then of course one-boxing is no longer rational since there’s no longer a reliable connection between your choice and the second box’s contents.

But there’s a deeper reason why the opacity matters, and it concerns the agent rather than the predictor.

With opaque boxes, the agent’s way of coming to know what’s in the box is practical. You can’t peek inside. The only way to arrive at a justified expectation about the box contents is to work out what you ought to do and recognise that the predictor, being sensitive to the very rational grounds you’re exercising, will have anticipated your conclusion. Your deliberation about what to do and your knowledge of what the predictor did are not two separate cognitive acts: the first is your route to the second just like your deciding to go to a party is how you know that you will be among the guests. That’s why your choice bears on the box contents: not by reaching back and changing them, but because the deliberation that produces the choice stems from the same rational ground that the predictor tracked.

With transparent boxes, all of this collapses. You look and see $1,000,000 or you look and see nothing. Your knowledge of the box contents is now derived from observation, not from practical deliberation. And once you’ve seen what’s there, your choice is decoupled from the box contents. And for a spectator, dominance reasoning is perfectly sound. Of course you take both boxes.

So yes, with transparent boxes I would two-box. But this doesn’t tell you anything about the opaque case. The transparent case converts you from a deliberator into a spectator. The opaque case preserves your status as a deliberator whose practical reasoning is the very thing the predictor was sensitive to. The question is whether being a deliberator in the opaque case gives you genuine causal reach over the box contents. I think my android delegate argument shows that it does.

The one-boxer isn’t someone who has fooled themselves into thinking it’s rational to leave money on the table. They’re someone who recognises that their practical deliberation, in circumstances where the predictor is sensitive to it, is the means by which the large prize is secured just like the android programmer’s design of the decision procedure is the means by which the large prize is secured. Nobody would say the programmer “fooled themselves” into programming the android to walk away from the $1,000. And the agent who one-boxes on rational grounds is simply doing directly what the programmer did indirectly.

Right, you are still assuming your determinism, you still don’t even understand what I mean by agency.

For the dominance argument to work, you need the agent to be powerless to determine the contents of the opaque box (that’s already settled by their prior dispositions as the predictor tracked them). But you simultaneously need the agent to be powerful enough to secure an extra $1,000 by deciding to two-box, which means deciding to be something other than what their dispositions (as tracked by the predictor) already determined them to be.

Yes, assume the agent is powerful like that, call that ghosts if you want, should the agent one box or two box?

Nothing substantial actually changes in the transparent case. The predictor can still make the prediction.

This seems to be completely fallacious. OK, given a useless definition of ‘power over’ to mean ‘ability to make something not what it is’, I suppose the argument works, but the conclusion is pointless.

The argument uses identical wording for two completely different meanings (unless we use the silly definition), and then equivocates due to the similar language.
First point seems to imply that one cannot have a causal influence on past events, but that implication requires a more causal definition of ‘power over’.
Second point, using the same wording, implies that one cannot have a causal influence on future events, which is just plain wrong. Certain future events are a function of choices you make now, so ‘power over’ must mean something else here. It is this causal influence that makes one an agent and thus responsible for consequences of said choices. Rocks don’t have this agency since they don’t have options to choose from.

Determinism has nothing to do with it. It is also true if randomness is the case instead of determinism. The argument also asserts determinism, something not known.

As for the Newcomb scenario, it rewards logical stability and decision-rule coherence. It does not in any way reward spontenaity, or metaphysical freedom.

I’ve brought much of this up before, and received no pushback on any of the points. I presume you concede them.

I think it comes from the definition ‘ability to have done otherwise’, normally pretty meaningless since one cannot alter the past, but this scenario rewards exactly that: ability to do other than what was predicted, and not only that (since that by itself can be done half the time), but to do it in such a way to increase utility over the rational method, and not even free will grants that.

There does very much seem to be a most rational action (regardless of your opinion which), and given that the consensus on this issue is so split, the implication is that humans are hardly perfectly rational agents.
To illustrate:

That’s an example of begging. It ‘proves’ that A=A by first presuming that A=A. This illustrates my point of at least some humans not being perfectly rational agents. Perhaps you can argue that determinism made you say that, and your would have provided a valid example if physics hadn’t forced you to do otherwise.

@Michael:
Why does the predictor fail occasionally? Perhaps it’s due to some deciding that unpredictability is the most rational course, which it isn’t.

No contest still. There’s 3 cards on the table, one of which is 9 of clubs, which wins 1M. I can choose one for free, or spend 1K to choose two of them. No looking of course until choice is made.

TB. There’s no rational capacity I’m required to exercise at choice time, even if the presence or absence of the 1M is based on that capacity.
Edit: I quickly take that back. OB. See post below. Kudos for fooling me for a minute.

Despite my initial reply, this is a cool example. The rules are not completely stated, so let’s keep as many of them the same as possible.

Both boxes are transparent. That is the sole difference. The predictor still makes his prediction with 99% accuracy and fills the 2nd box accordingly. What do you choose? 1B then. Yea, that sounds pretty irrational, but this still is an exercise of my rational capacity.

I find it fascinating that the opaqueness of the 2nd box is not critical to the scenario. Thanks for that suggestion.

I don’t understand your criticism of the Consequence argument but it doesn’t matter, I am not making the exact same argument, just a similar one in spirit.

Rocks don’t have this agency since they don’t have options to choose from.

They don’t have options to choose from and agency because we assume rocks are deterministic entities. It follows from determinism.

As for the Newcomb scenario, it rewards logical stability and decision-rule coherence. It does not in any way reward spontenaity, or metaphysical freedom.

I’ve brought much of this up before, and received no pushback on any of the points. I presume you concede them.

That it rewards logical stability and decision-rule coherence? I am not sure what that means but I can grant you the point.

That’s and example of begging. It ‘proves’ that A=A by first presuming that A=A. This illustrates my point of at least some humans not being perfectly rational agents

You’re not even engaging with what I am saying. I am literally acknowledging that it could be considered begging the question, I then show my ‘syllogism’ can still be a valid one with an example. If the argument is valid, it could be that Michael disagrees with some premises, I ask if they do disagree.

There is no issue whatsoever. If I say “Assume A, then A” I have shown A \implies A and that’s perfectly valid and true. That’s jut not what begging the question is.

It’s an interesting question, because with other problems like this — e.g. Sleeping Beauty and Two Envelopes — we assume that the agent is perfectly rational, but a perfectly rational agent can only take the most rational course of action, and so the notion that the prediction is only near-perfect is problematic; it suggests that in some cases it predicts that the perfectly rational agent will take the least rational course of action, which makes no sense.

So strictly speaking it can only fail if the agent either is not perfectly rational or is predicted to not be perfectly rational.

This is where the distinction between the perfect prediction and the near-perfect prediction is important. Some claim that if the prediction is perfect then one-boxing is the most rational course of action, but that if the prediction is only near-perfect then two-boxing is the most rational course of action. Whereas I say that the prediction will be perfect if it is known that everyone will take the most rational course of action (one-boxing) and that the prediction will be imperfect only if the agent either takes the least rational course of action (two-boxing) or is predicted to take the least rational course of action (two-boxing).

After all, it seems like special pleading to say that taking all of the money is more rational than taking some of the money, except when the prediction is perfect. If someone accepts that there is at least one occasion where taking some of the money is more rational than taking all of the money then they need to explain (without begging the question) why the prediction being near-perfect isn’t also one of these occasions — or accept that the prediction being near-perfect doesn’t make a difference.

I think the opaqueness is more critical than you suggest. Making the boxes transparent doesn’t just give the agent more information. It also changes the structure of the problem in a way that more closely mirrors Kavka’s toxin puzzle.

When the second box is opaque, you lack the power to deliberately act in a way that renders the prediction false. Your deliberation about what to do is your route to the outcome, and the predictor’s sensitivity to that deliberation is what makes one-boxing effective. But when the second box is transparent, an agent who is naturally inclined to one-box and who sees $1,000,000 in the box now has the power to two-box (in spite of their prior inclination to one-box) knowing it won’t penalize them. The initial inclination to one-box, and the consequent action from the predictor, already did its job. The rationale for actualizing their one-boxing disposition lapses just as, in Kavka’s puzzle, the rationale for drinking the toxin lapses once the prize has been secured.

And the Kavka parallel runs deeper. It isn’t just that the agent can defect upon seeing the million — it’s that their foreknowledge of this undermines the intention upstream. Just as Kavka’s agent finds it hard to intend to drink the toxin because they can foresee that their future practical situation won’t sustain it, the transparent-box agent’s resolve to one-box is undermined by their knowledge that they’ll face a situation where two-boxing costs them nothing. You might walk into the room with the resolve to one-box but that resolve has the same fragile, self-undermining character as the Kavka agent’s resolve to drink.

I’ll have more to say in a follow-up post regarding the question why this means the agent’s ignorance in the standard opaque setup isn’t a blindfold that enables self-deception but something structurally empowering precisely because of what it enables the predictor to do (which is, to anticipate, not just to reward or punish the agent’s predispositions but their actual exercises of those predispositions).