A branching treasure-chest task reveals how researchers tested planning under changing rewards and unexpected outcomes.

Study: Environmental stochasticity reduces human planning effort
A recent study published in the journal Nature Communications suggests that people reduce planning effort when rewards become less reliable, change more frequently, or depend less consistently on their chosen actions.
In three separate online experiments, participants responded more quickly before their first choice as environmental randomness increased. The findings are consistent with people weighing the cognitive costs of planning against the potential benefits when deciding how much effort to devote to planning.
Stochasticity, or randomness, in real-world situations can make planning more challenging. Stochasticity is common in real-world settings, but how people balance the cognitive costs of planning in such situations remains unclear. Random factors, such as unexpected disease pandemics or shifting climate patterns, are often beyond individual control and can disrupt carefully made plans.
In situations with imperfect information or outdated recommendations, the environment may be considered unreliable, and the reliability of information can modulate cognitive effort. In some situations, the reward is known but can vary unpredictably over time, a phenomenon known as volatility. Uncertainty in the outcomes of planned actions can also influence decision-making. In stochastic environments, determining the best plan may be cognitively demanding.
About the study
In the present study, participants completed planning tasks that required them to navigate a branching decision tree and choose routes that could yield more points under different levels of volatility, reliability, and controllability. Each experiment included 100 participants (50 male, 48 female, and two nonbinary), with mean ages of 36.9, 37.8, and 38.0 years for the reliability, volatility, and controllability experiments, respectively. All participants were adults, fluent in English, and residing in the United States (US). Across the experiments, 300 unique participants each completed one task variant.
To explore how reliability influences human planning, the researchers asked participants to complete a task remotely. They navigated through several treasure chests labeled with values from one to nine and tried to earn as many points as possible. At each branch, participants chose whether to move left or right, and that choice determined which chest they encountered next. To manipulate reliability, mystery chests whose numbered labels did not predict their randomly assigned rewards were also included. Participants learned which chests were mystery chests only after selecting them. In exploratory analyses, the team considered first-choice response times as an indirect measure of planning effort.
To explore environmental volatility, the researchers introduced a chance that chests in later rows would be resampled at each step. Modeling the uncertainty of distant outcomes, chests located further away were more likely to change before participants reached them. The team also explored the impact of transition noise, which reflected controllability, on planning effort. To do so, they varied the probability of reversing the selected direction from 0% to 50%, so that choosing left could result in a rightward movement, and vice versa. Participants were informed of the level of stochasticity in each task block. The researchers used generalized linear mixed-effects models (GLMMs) and linear mixed-effects models (LMMs) to estimate effects on choices, earnings, and response times.
To examine planning and the processes underlying participants’ decisions, the researchers developed several computational cognitive models to compare possible decision strategies across the tasks.
The models were used to estimate which options participants would choose at each stage. The researchers assumed that participants compared the perceived value of the left and right branches when making each choice. They also compared alternative models in which participants considered a limited number of future steps, chests above a value threshold, or the highest-ranked chests.

Each game consists of a triangular arrangement of treasure chests. At each step, participants choose either left or right as they progress down the tree (the arrows here are shown only for illustration). B Reliability: with probability q, the participant encounters a mystery treasure chest, which returns a random value between 1 and 9. C Volatility: with probability q, any treasure chest in the rows below the participant’s current position is redrawn with a new random value between 1 and 9. D Controllability: with probability q, the participant’s choice (left or right) is flipped.
Results
The same trend emerged across the three types of stochasticity. As conditions became more stochastic, participants responded more quickly before their first choice, consistent with reduced planning effort. The best-fitting models also indicated that participants became less sensitive to differences between the estimated rewards of alternative routes as stochasticity increased. This pattern was consistent with policy compression, meaning simpler, less precise decision policies.
Model comparisons suggested that, rather than calculating the expected value of each option, participants used simpler strategies that treated outcomes as certain, a process called determinizing. The numbers shown on the chests also affected participants’ choices. They tended to choose the chest with the higher displayed value, especially when the two options differed considerably, but this preference weakened as stochasticity increased. The expected-value model was excluded from the volatility experiment because its calculations were computationally intractable.
Participants consistently earned more points under low-volatility conditions than under high-volatility conditions. Likewise, participants scored significantly higher when they had greater control over the outcome or more reliable reward information. The best-performing models used depth filtering, with planning depth fixed across stochasticity levels while sensitivity to value differences changed. The strongest modeling evidence concerned changes in decision precision.
Conclusion
Overall, the findings suggest that participants use simpler decision policies and may reduce cognitive effort through policy compression in stochastic environments. The tasks examined irreducible randomness. The authors suggest that uncertainty people can reduce through learning may encourage greater planning effort. Applicability to everyday planning remains untested.
In future studies, researchers should explore the influence of stochastic environments on cognitive development and future-oriented thinking.
Journal reference:
- Lei, J., Olieslagers, J., Arfaei, N., Lin, D. X., & Ma, W. J. (2026). Environmental stochasticity reduces human planning effort. Nature Communications. DOI: 10.1038/s41467-026-78023-9, https://www.nature.com/articles/s41467-026-78023-9