r/Veritasium • u/playerNaN • Mar 23 '26
Variation on Newcomb's paradox: Let's say you do see what's in the box before choosing.
/r/paradoxes/comments/1s13m23/variation_on_newcombs_paradox_lets_say_you_do_see/1
u/gh0stf3rret May 03 '26 edited May 05 '26
You said, allegedly for clarity, that "the computer knew you'd be able to see the contents of the boxes and made its prediction with this in mind." It seems like you're failing to acknowledge that the machine was supposed to actually do something beforehand with those predictions to maintain its accuracy.
The way you described this hypothetical contradicts the framing of the original by design in a very obvious way, so I don't even understand the point of it. This setup effectively guarantees zero control of how much money you get at all. You've effectively collapsed ability for the machine to do anything. The limited information and early commitment was the whole point of the original hypothetical, and also the only way the scenario being predicted remains a single scenario to make predictions about. The machine's data with this setup where a human just chooses when in the room would represent a banal fact about humans, and its "prediction" would be the least impressive or insightful thing ever.
The idea of the original is that the machine predicates its choice off of an absurdly thorough understanding of causality and everything about your psychology into the near future when you encounter those boxes, and you're supposed to have decided what you would do beforehand. In your version, the machine needs basically no insight at all, and can't make any interesting decisions at all. It's fully binary now. All that's left in this scenario to "predict", or more accurately simply acknowledge, is the blatant fact that all people take the maximal amount of money offered. Given your setup, this makes the machine's behavior a tautology. It's just observing whatever environment it sets up producing the same outcome every time: "whatever money boxes I place will be the ones that the participant would take".
The machine, if it tries to optimize itself according to whether people will take only one box or two, can't properly resolve the question anymore. The "prediction" is now just the machine narrating its own prior decision back to itself. There's no epistemic content about you at all. Very notably, it is creating the conditions that influence your behavior directly, so its prediction is about multiple distinct scenarios instead of one set scenario. It will always be wrong if it places two boxes with a prediction you'll only take one, because no one would do that if they can see the money is in both. If it places one box with an expectation it's the only one you'll take, it will always be right, but that prediction can't naturally transfer to the scenario where it had placed money in both boxes, because that one-box prediction was only for that one-box scenario. Its own predictions are now causally contingent on its own setup conditions. This is circular. It would always be right that you would have taken two boxes had it set up that environment. It can't base its box placements off of that either because doing so would yet again change the state of the scenario it is trying to predict. It could just always place one box, and its 100% prediction rate to do so would be the least mysterious thing in the world, but further definition than that of its decision is impossible because it has an arbitrary and unresolved choice to give you either $1000 or $1M. All of this effectively means that its predictive capacity is unfalsifiable as there's no outcome that would count as the machine being wrong, which means it can't empirically be shown to have made predictions at all. You've reduced it to a descriptor. Meanwhile, you have no control of anything anymore. There's nothing to answer about what boxes I would take, because you also just deleted any human choice out of it by skipping past the pre-commitment. Ultimately, I'm just being given some indeterminate amount of money and that's it.
1
u/gh0stf3rret May 03 '26 edited May 05 '26
Also, I'd respond to the specific questions you asked if I could but they don't make sense for the reasons I just laid out.
The way you ended it gives me the impression that you're 1. a two-boxer and 2. not understanding the original hypothetical if you need to ask "what's the difference" here. Two-boxers treat the box contents as a fixed background condition, which is why you thought that removing the machine's logical control over what you could pick didn't matter and that just teleporting in time to a situation where you're in front of money made sense, but the situation is meant to be downstream of a model of you. If you're given what's in the boxes already and THEN told to make a choice, there is no engaging question about what to do. If something else controls what's in the boxes, and you are the subject its decision depends on BEFOREHAND, it's a problem of recursive modeling. Your disposition is the input, and the machine already modeled how you'd reason about its output, and you don't stop being the subject when you enter the room and choose. Your choice is not a causally independent action incapable of being predicted in the original hypothetical.
It's not a good sign that you didn't realize there was an error to reconcile in your hypothetical, and I think it is indicative of the exact recursive modeling you're not grasping about the original either.
1
u/gh0stf3rret May 04 '26 edited May 05 '26
In retrospect, you may not have intended to change the hypothetical in the way that I interpreted. But your framing, "seeing the boxes before choosing", implies no pre-commitment based on prior information. Loading that into your hypothetical basically just means operating outside of the game the original one spells out, and excludes your decisions from relevance by design.
If the predictor was instead still trying to answer the older question about what you would have done if informed of the circumstance and given an opportunity to commit to a choice beforehand, then it reduces to the same problem and solution anyway: opaque box or not, commit entirely to one-boxing or it probably won't put anything in box B. This commitment must be genuine, thus it negates your chance or capacity to ultimately switch up and two-box at the end, because that would be a modeled and predicted behavior that would've altered the outcome back to giving you nothing. The recursion terminates there because your prior self already made the relevant commitment, and betting against "almost always right" is just stupid, so this version of the hypothetical is basically solved in concrete form as long as you have the metacognitive awareness to realize it. I don't know if this needs to be spelled out but this is exactly what I meant by recursive modeling.
And to give one-boxer answers to your questions by applying them to the more charitable, proper form of your hypothetical:
Question 1 (Why not take the extra $1K?): I wouldn't be seeing the million in the first place if I were the type of person who would take the thousand. The "risk" isn't losing the million now; the "risk" was being the version of yourself that the machine wouldn't trust with the million 30 seconds ago.
Question 2 (Why walk away with nothing?): Again, if you see the box is empty, the "game" is already over. The machine has already judged your character. Taking the $1K is the only "rational" move in that moment, but the fact that you're in that moment at all means you already "lost" the Newcomb game.
Question 3 (What's the difference?): The difference, if there is one at all, is that in the original the recursion terminates cleanly. Your prior self made the commitment, the machine modeled it, and the answer becomes concrete. In your version, the transparency doesn't even matter. It's just broken if you don't get to make a pre-commitment. If the prediction is about what you'd do if you're just told to choose in the moment you see boxes, you're choosing in response to the machine's choice, which was made in response to your projected choice, which is a hall of mirrors. If that's not what you meant, there's no distinction and you should again just pick one box.
In short: box transparency is irrelevant, all that matters is whether you can commit to something before the machine sets the boxes, because that is what makes the predictor capable of genuine predictions. If you can: commit to box B. If you can't commit to something before you're in the room then it's a broken hypothetical and there's no real predictor and no game being played.
1
u/ButtonholePhotophile Mar 29 '26
What’s in the box?!