Reasoning Machines and Mediational Theory

96

Guest Post by Hank Schlinger, Eb Blakely, and Jiawei Xiang

In our previous post, we described a mediational theory of problem-solving. In one respect, it is not really a theory of problem-solving as much as a description of how we verbal humans solve problems most of the time. In simple terms, we react to stimuli (the problem) by engaging in a cascade of verbal responses that may or may not result in solving the problem. Sometimes the verbal responses are relatively simple as, for example, in the case of a three-member equivalence class, in which a child is asked to point to a picture of a cat. When she hears “cat” or “point to the cat,” she echoes “cat,” the form that is reinforced at the same time the selection response is either explicitly by the teacher or researcher or automatically. As a result, when shown a picture of a cat and asked what it is, she can now say “cat.” In this example, she reacts to “point to the cat” by echoing “cat,” which when reinforced in the presence of seeing the picture, converts the echoic response form into a tact. Problem solved. No need for multiple exemplar training or complicated jargon-laden theories.

The analysis is similar for instances of strictly verbal stimuli. For example, suppose a child is shown an unfamiliar object like a record album and is told “this is a record album,” and as a result can subsequently say “record album” when shown the object and asked “What is this?” Again, when told “this is a record album,” he echoes “record album,” either out loud or silently, in its presence. If he echoes out loud, the speaker is likely to say something like “right, it is a record album.” If he echoes silently, the match between what he “hears” himself say and what he heard the speaker say automatically reinforces his echoic response. Either way, the form of the echoic response is reinforced in the presence of the record album, thus converting it into a tact. Still, no need for multiple exemplar training. In a sufficiently verbal individual, one trial is enough.

We then described some implications of a mediational theory for clinical practice, for example, the use of differential observing responses and instructive feedback. We suggested that “it is unnecessary in most instances to provide multiple exemplar training, which sometimes can involve dozens if not hundreds of trials. Instead, a mediational theory suggests that practitioners can target directly the mediating responses necessary for the conditioning of the desired relations. In other words, as a practitioner you would be teaching problem-solving responses so that you would not have to rely on chance, luck, or accidental conditioning resulting from a great many multiple exemplar trials.”

We concluded the blog post by suggesting that a mediational theory of behavior has theoretical implications for what are called reasoning large language models (LLMs) of artificial intelligence (AI). In general, when an LLM, or a human, is asked to solve a problem, there are mediating steps that, if correct, will end up producing the correct answer. Humans learn these steps and how to apply them in a variety of situations. Likewise, LLMs must “learn” them as well. Moreover, there is a symbiotic relationship that we see between a mediational theory and LLMs, where both demonstrate that mediating steps are crucial in solving problems and in reasoning.

Traditional LLMs of AI have historically lacked reasoning capabilities. They operated by rapidly predicting subsequent words based on statistical probabilities, provided that all necessary information was explicitly given. According to one AI researcher, these models are typically “trained to give the answer in one quick shot,” which is not how humans solve problems. There is, however, another kind of model called reinforcement learning with verifiable rewards (RLVR) that involves intermediate, or mediating, steps that lead to the final answer.  

In the terminal reward model, correct answers to a problem increase the probability of each mediating step. However, this model doesn’t consider the accuracy of each step. Thus, faulty logic is strengthened, as along as the correct answer is produced. This issue is addressed in the process reward model. In this model, each mediating step is evaluated for accuracy, and probabilities are adjusted accordingly. For an incorrect step, another candidate step can be selected and evaluated. If the correct answer is produced after all of the mediating steps are completed and the steps are all correct, this adds to the strength of each step. But let’s say that a mediating step is incorrect, but a correct answer is still produced. In this case, there is no terminal reward, or strengthening of mediating steps, because the logic was faulty.

Note that these two models—the terminal reward and process reward—simulate real-life teaching strategies. An algebra teacher may only grade the students’ final answers, irrespective of how they were obtained (i.e., terminal reward). Or the teacher might require students to show their work in addition to the final answer, and the grade depends on obtaining the correct answer, but also providing correct, logical steps to obtaining the answer (i.e., process reward). The process reward method is the most effective way of solving problems for humans, as well as for computers. The emphasis is on both obtaining the correct answer and the mediating steps that lead to the answer.

The development of these more advanced LLMs is important as they more closely align with human problem-solving and reasoning. When people confront even a trivial problem, they do not typically produce a final answer in one uninterrupted emission. Instead, they immediately engage in a sequence of mediating verbal (and imaginal) responses, breaking problems into sequential steps either overtly or covertly. Each of those steps is evoked and followed by stimuli, which function both as conditioned reinforcement for the preceding step and discriminative for the next step. Large language reasoning models explicitly replicate this human-like mediation process. As explained by another AI researcher, a reasoning model “takes time to break a question down into individual steps and works through a ‘chain of thought’ process to arrive at a more accurate answer.”

For instance, one researcher provided an interesting example as follows: “A juggler can juggle 16 balls. Half of the balls are golf balls, and half of the golf balls are blue. How many blue golf balls are there?” In a reasoning model, rather than generating an immediate final answer, the system works through the problem step by step just as a human would do:

  1. There are 16 balls in total.
  2. Half the balls are golf balls.
  3. That means that there are 8 golf balls.
  4. Half of the golf balls are blue.
  5. That means that there are 4 blue golf balls.

As this researcher explains, this step-by-step verbal production “encourages the model to generate intermediate reasoning steps rather than jump directly to the final answer, which can often (but not always) lead to more accurate results on more complex problems.”

OpenAI similarly describes reasoning models as systems that “think before they answer, producing a long internal chain of thought,” enabling them to excel at tasks requiring complex, multi-step cognitive performances such as scientific problem-solving and programming. This deliberate, stepwise processing aligns closely with a mediational theory of human problem-solving. In fact, behavior analyst and programmer Bill Hutchison has created computer simulations that have demonstrated the sufficiency of a behavior-analytic mediational analysis, particularly the role of conditioned reinforcement of each intermediate step.

To conclude, the evolution of AI reasoning systems toward explicit step-by-step verbal mediation, established and maintained by differential consequences, strongly supports core aspects of a mediational theory of verbal relations. These technological developments, arrived at entirely without reference to behavioral theory, underscore the explanatory power of mediating behaviors established by reinforcement and align closely with theoretical propositions that behavioral science has advanced for decades.

Conversely, a mediational theory provides an account of how mediating responses emerge, what determines their form, and how terminal reinforcement modifies the evocative function of each preceding contextual stimulus so that every link in a behavioral chain becomes more or less probable as a function of the reinforcement history that chain has accumulated. It is precisely this account, grounded in the basic principles of operant conditioning and requiring no new terms, principles, or theoretical constructs, that a mediational theory can contribute to current AI research. A behavior-analytic mediational theory of verbal relations offers AI researchers a scientific framework that goes beyond empirical trial-and-error, one that may inform the design of training procedures capable of producing more robust, flexible, and human-like reasoning in artificial systems.

Leave a Comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.