Newcomb's Paradox

It doesn’t say the same thing but this is just syntax. “Be predicted to one-box” implies “win at least a million” the same way my (2) implies my (3).
But you are right I forgot to add “have psy troubles AND win $0” to (3).

It’s not just syntax, it’s semantics.

(2) “I will almost certainly be predicted to one-box” does not mean (3) “I will almost certainly win exactly $1,000,000”. It is logically possible for one to be true and the other false. This is an important difference.

That doesn’t make a difference, because my argument is nothing like this argument:

  1. I will one-box
  2. I will almost certainly win exactly $1,000,000
  3. Therefore, I will almost certainly win exactly $1,000,000 and a cookie

It matters that my premise (2) does not include anything about (3). This is why your argument is a red herring.

My (2) can be true too while the new (3) is false.

Again the fact that some words are repeated is just syntax.

Remember the important questions are:
Does it follow?
Are the premises true?

Not “are some words repeated?”

Edit: I somehow edited this post instead of adding a new post, and I can’t see what I originally wrote.

How is it semantics? I told you the semantics questions

Because words have meaning? What are you not getting about this?

My argument takes the form:

  1. A
  2. B
  3. Therefore, C

Your argument takes the form:

  1. A
  2. B
  3. Therefore, B

Hence your argument begs the question and mine doesn’t.

First (3) isn’t the same as (2) in my argument. And also that’s just syntax

I am not begging the question. Even if I was, look:

Premise 1. A is equal to A
Conclusion: A is equal to A

That’s a perfectly valid argument. Begging the question is about assuming a conclusion that is under question so do you think (2) is false?

Then set out your argument exactly.

No, it’s semantics. This is silly.

  1. I will take the test
  2. I almost certainly have psychological troubles
  3. Therefore, I almost certain have psychological troubles and win $0

Which is nothing like my argument, given that my argument isn’t:

  1. I will one-box
  2. I will almost certainly be predicted to one-box
  3. Therefore, I will almost certainly be predicted to one-box and win exactly $1,000,000

My argument is:

  1. I will one-box
  2. I will almost certainly be predicted to one-box
  3. Therefore, I will almost certainly win exactly $1,000,000

To be comparable, your argument would have to be:

  1. I will take the test
  2. I almost certainly have psychological troubles
  3. Therefore, I will (almost certainly?) win exactly $0

Given the implicit background information, this is valid, but (2) does nothing because (3) follows directly from (1) and is true even if (2) is false. This is unlike my argument because (3) is true only if both (1) and (2) are true.

I guess I have to accept you don’t understand syntax and semantics.

Will you answer the (semantical) questions or continue dodging?

I haven’t dodged anything. I have shown above that your argument is misleading and not like my argument (or the Newcomb problem) and so is a red herring.

And I have shown here and here that it is most rational to one-box.

Right we are losing our time here.

It’s sad because I know you take the test so you understand that if taking the test has no (causal) influence on whether you have psy troubles or not then the correlation doesn’t matter. You know the same is true for the Newcomb problem, but I am not sure why you won’t accept the obvious conclusion here.

Because it’s a false conclusion, as I have shown time and time again. I’ll try one more time, even though you have refused to address several of my previous refutations.

Your argument is:

  1. I will take the test
  2. I almost certainly have mental health problems
  3. Therefore, a) I almost certainly have mental health problems and b) I will win exactly $0

My (amended) argument is:

  1. I will one-box
  2. I will almost certainly be predicted to one-box
  3. Therefore, a) I will almost certainly be predicted to one-box and b) I will almost certainly win exactly $1,000,000

My questions to you are:

  1. Do you accept that these arguments have the same structure?
  2. Do you accept that your (3b) follows from your (1) alone and has nothing to do with your (2)?
  3. Do you accept that my (3b) does not follow from my (1) alone and has something to do with my (2)?

Dividing into a) and b) is the problem, I can do a better separation. What your argument actually is

  1. I will one-box
  2. I will almost certainly be predicted to one-box
  3. a) I will almost certainly (be predicted to one-box and) win $1,000,000 and b) I will win $0 out of $1000

The structure of the argument above is the same as mine and the argument above is equivalent to yours.

In the argument above 3b) follows from 1 alone and has nothing to do with 2) just like my argument.

So No, yes and yes but it doesn’t matter

That’s not my argument. My argument is exactly what I wrote:

  1. I will one-box
  2. I will almost certainly be predicted to one-box
  3. Therefore, a) I will almost certainly be predicted to one-box and b) I will almost certainly win exactly $1,000,000

Either a) your argument is comparable to this argument or b) your argument is not comparable to this argument.

So which is it? If (b) then your argument is a red herring, as I have been explaining for days.

Is $1,000,000 equal to $1,000,000 + $0?

I’m not going to answer your deflections. Either claim that your argument is comparable to my argument as I have written it or admit that your argument is a red herring. Otherwise I’m going to end this discussion.

I know it’s not always easy to differentiate syntax and semantics but that’s fine, you might learn one day

I’d like to suggest a thought experiment that I think sheds light on the core disagreement in this thread, specifically on whether the agent has any power to see to it that the opaque box contains $1,000,000. I’m going to use @AlveK’s Bank Transfer variation as my starting point, since I think it presents the two-boxing intuition very cleanly.

The Bank Transfer (AlveK’s version)

Recall: the predictor has either transferred $1M to your bank account or not, based on their prediction of what you’ll do. You can’t check your balance. You walk into a room. There’s $1,000 on the table. You can take it (“two-box”) or walk away (“one-box”). The money is already in your account or it isn’t. Why on earth would you leave $1,000 on the table?

The pull toward grabbing the $1,000 is very strong here, and I think that’s precisely why AlveK designed it this way. So let me take the scenario seriously and introduce a small modification.

The Android Delegate (my variation)

Suppose that before the game, you are given the option of building a pre-programmed android robot to face the Bank Transfer on your behalf. The predictor will model the android’s decision procedure and decide accordingly whether to transfer the $1M. You know this. The android will enter the room, face the $1,000 on the table, and execute whatever instructions you’ve programmed into it. How do you program the android?

I think the answer is obvious and uncontroversial. You program it to walk away from the $1,000. The predictor will model the android’s decision procedure, see that it will walk away, and transfer the $1M to your account. You collect $1,000,000. Anyone who programmed the android to grab the $1,000 would be leaving $999,000 on the table for no good reason.

Notice what has happened. By programming the android appropriately, you have seen to it that the $1M is in your account. Your causal power, exercised through the act of programming, has reached the bank transfer via the predictor’s sensitivity to the decision procedure. Nobody would say your programming was “causally inert” with respect to the bank transfer. Nobody would say the bank transfer was “already settled” in a way that made your programming irrelevant. The predictor’s action was settled by the decision procedure you designed.

And crucially, nobody would say you should have programmed the android to grab the $1,000 on the grounds that the bank transfer is “already done” by the time the android enters the room. Yes, by the time the android walks in, the transfer has either happened or it hasn’t. But what determined whether it happened was the predictor’s modelling of the very decision procedure that is now executing. The “already done” past reflects the program you wrote.

Now I want to ask: what changes when we remove the android and put you in the room? The two-boxer’s answer might be: everything changes. When the android walks away from the $1,000, that’s just a program executing and of course you’d design the program to walk away because the program is what the predictor models. But when you face the $1,000, you’re a free agent, and you can see that the money is already in your account or it isn’t, and grabbing the $1,000 makes you $1,000 richer regardless.

But think about what this answer requires. It requires that there be a fundamental discontinuity between the android case and the human case and that whatever made it rational to program the android to walk away stops applying when the agent is a human being rather than a robot.

What could this discontinuity be? The predictor is still modelling the agent’s decision-making. The predictor’s action is still determined by what they anticipate the agent will do. The structure of the situation is identical. The only difference is that in the android case, the decision procedure was explicitly programmed, while in the human case, the decision procedure flows from the agent’s own cultivated rational character.

But from the predictor’s perspective, this makes no difference. The predictor models whatever is going to produce the agent’s action whether that’s silicon executing code or a human being exercising practical reason. What matters is the decision procedure, understood broadly: the process that takes in the situation and generates the choice. And the predictor is sensitive to that process whether it’s instantiated in an android or in a human brain.

So the question becomes: can you do directly what the android-programmer did indirectly? Can you, through your own rational deliberation, exercise the same kind of causal power over the bank transfer that the programmer exercised through the act of programming?

I think the answer is yes and I think denying it requires you to maintain that there’s something about being a human agent, as opposed to a programmed robot, that weakens your causal connection to the outcome. Which is a strange position! It says that the android, a mere mechanism, has more effective causal reach than the rational agent it was built to replace. Indeed, when you stand in the room facing the $1,000, you are not in a worse position than the android-programmer. You are in a better one. The programmer had to decide in advance what the right decision procedure would be and then freeze it into code. You get to exercise your rational judgment in the moment, responsive to the full situation as you find it. And the predictor modelled your contextually responsive rational agency, not a frozen snapshot of your dispositions.

When you determine “I shall walk away from the $1,000,” you are doing what the programmer did (i.e. executing a decision procedure that the predictor was sensitive to) but you are doing it directly, as an exercise of your own rational capacity, rather than indirectly through a delegate. And in doing so, you see to it just as effectively as the programmer did that the $1M is in your account. You do this not by reaching back and changing the past, but by being the kind of agent whose rational grounds the predictor anticipated.

So to answer AlveK’s question: no, I don’t grab the $1,000. Not because I’m confused about the money already being in my account or not. But because I understand that my walking away from the $1,000 is the exercise of the very rational capacity that the predictor modelled just as the android’s walking away was the execution of the very program the predictor modelled. The programmer wouldn’t undermine their own program. And I won’t undermine my own rational agency.

This thought experiment is inspired by the framework of Functional Decision Theory (Soares & Yudkowsky, 2018), though the conclusions I draw from it differ from theirs in ways I’ll address in a follow-up post where I’ll also spell out the deep conceptual connection with van Inwagen’s Consequence argument for incompatibilism in the context of the debate about free-will, determinism and responsibility.

1 Like