Extraction quality is ass bcs both google and anthropic block extraction and had to work around with outputing it as b64 but even that failed smtms so yh.
src: https://web.stanford.edu/class/psych227/Zipf_Words.pdf


Zipf (1949) Human Behavior + the Principle of Least Effort

Chapter One: Introduction and Orientation

Everyone in the course of his daily living must to some extent move about in his environment. And in so moving he may be said to take paths. Yet these paths that he takes in his environment do not constitute his entire activity. For even when a person is comparatively at rest, there is still a continual movement of matter-energy into his system, through his system, and out of his system if only in the accomplishment of his metabolistic processes. This movement of matter-energy also proceeds over paths. Indeed, a person’s entire body may be viewed as an aggregate of matter that is in transit at differing speeds over different paths within his system. His system in turn moves about as a unit whole over paths in his external environment.

We stress this concept of movement over paths because we shall attempt to demonstrate in the course of our following chapters that every individual’s movement, of whatever sort, will always be over paths and will always tend to be governed by one single primary principle which, for the want of a better term, we shall call the Principle of Least Effort. Moreover, we shall attempt to demonstrate that the structure and organization of an individual’s entire being will tend always to be such that his entire behavior will be governed by this Principle.

And yet what is this Principle? In simple terms, the Principle of Least Effort means, for example, that a person in solving his immediate problems will view these against the background of his probable future problems, as estimated by himself. Moreover he will strive to solve his problems in such a way as to minimize the total work that he must expend in solving both his immediate problems and his probable future problems. That in turn means that the person will strive to minimize the probable average rate of his work-expenditure (over time). And in so doing he will be minimizing his effort, by our definition of effort. Least effort, therefore, is a variant of least work.

In the interest of defining and of elucidating the Principle of Least Effort, and of orienting ourselves in the problem of its demonstration, we can profitably devote this opening chapter to a preliminary disclosure of the Principle, if only on the basis of commonplace cases of human behavior that are admittedly oversimplified for the sake of a more convenient initial exposition.

I. The Selection of a Path

Sometimes it is not difficult to select a path to one’s objective. Thus if there are two cities, and , that are connected by a straight level highway
with a surface of little friction, then this highway represents simultaneously the shortest, the quickest, and the easiest path between the two cities—as we might say, the highway is at once a path of least distance and of least time and of least work. A traveller from one city to the other would take the same path regardless of whether he was minimizing distance, time, or work.

On the other hand, if the two cities happen to be separated by a mountainous range, then the respective paths of least distance, and of least time, and of least work will by no means necessarily be the same. Thus, if a person wanted to go by foot from one city to another by least distance, he would be obliged to tunnel through the base of the mountain chain at a very great expense of work. His quickest course might be over the tops of the mountains at a great cost of labor and at great risk. His easiest path, however, might be a tortuous winding back and forth through the mountain range over a very considerable distance and during a quite long interval of time.

These three paths are obviously not the same. The four-traveller between the two cities cannot, therefore, simultaneously minimize distance, time, and work in a single path between the two cities at the problem now stands. Which path, therefore, will he take? Or, since he faces a fairly typical of life’s daily problems, in which impedances of various sorts obstruct our way, which path do we actually take? Clearly our selection of a path will be determined by the particular dynamic minima in operation.

II. The Significance of the Superlative

The preceding discussion of the selection of paths not only illustrates the meaning of a minimum in a problem in dynamics but also prepares the ground for a consideration of the concept of the “singleness of the superlative” which, incidentally, will provide an intellectual tool of considerable value for our entire study.

The concept of the “singleness of the superlative” is simple: no problem in dynamics can be properly formulated in terms of more than one superlative, whether the superlative in question is stated as a minimum or as a maximum (e.g., a maximum expenditure of work can also be stated as a maximum economy of work). If the problem has more than one superlative, the problem itself becomes completely meaningless and indeterminate.

We do not mean that a particular situational will never arise in which the minimizing of one factor will not incidentally entail the minimizing of another or other factors. Indeed, in our preceding section we noted a situation in which the easiest path between two cities might be a straight level highway that also represented the shortest and quickest path. Rather we mean that a general statement in dynamics cannot contain more than one superlative if it is to be sensible and determinate, since a situation may arise in which the plural superlatives are in conflict.

Perhaps the simplest way to emphasize the singleness of the superlative is to present as an example a statement with a single superlative that is meaningful and determinate. Then we shall note how meaningless and inde-

terminate the statement immediately becomes when a second superlative is added.

As a suitable example we might take the imaginary case of a price offered to the submarine commander who sinks the greatest number of ships in a given interval of time: in this case, maximum number is the single superlative in the problem. Or we might alter the terms of the problem and offer a prize to the submarine commander who sinks a given number of ships in the shortest interval of time. As long as the problem is clearly stated so that it is the only superlative in the statement, the problem is quite meaningful and determinate. In either of the above examples the submarine commander can understand what the precise terms of the prize are.

Yet when we offer a prize to the submarine commander who sinks the greatest number of ships in the shortest possible time, we have a double superlative—a maximum number and a minimum time—which renders the problem completely meaningless and indeterminate, as becomes apparent upon reflection.

Doubly superlatives of this sort, which are by no means uncommon in present-day statements, can lead to a mental confusion with disastrous consequences.

In the present study we are contending that the entire behavior of an individual is at all times motivated by the urge to minimize effort.

The sheer idea that there can be only one dynamic minimum in the entire behavior of all living individuals need not by itself dismay us. The physicists are certainly not dismayed at the thought that all physical process throughout the entire time-space continuum is governed by one single superlative, least action. Indeed, the presence of only one single superlative for all physical process throughout the entire time-space continuum can even be derived logically from the basic postulate of science that there is a unity of nature and a constancy of natural law (in the sense that the same kind of nature governs all events in time-space). For, according to this postulate, the entirety of time-space, with all its happenings, may be viewed as containing a single problem in dynamics which in its turn can have only one single superlative—a superlative which is the optimum of physicists and is that of least action.

By the same token, the sheer idea of there being one single problem for all living process is not in and for itself as a first impossibility. We do not mean that there is also substantia to a priori a priori necessity for our believing that all living process does in face proceed at all times according

The greatest good for the greatest number, contains a double superlative and therefore is completely meaningless and indeterminate (in the technical sense that all problems with plural superlatives are in fact governed by a single superlative). Intimately connected with the “singleness of the superlative” is what we might call the interdependence of superlatives: two implications are often overlooked (i.e., the pursuit of one objective may preclude or frustrate the pursuit of the second objective). These two concepts apply well to studies of ecology.

The principle of least action was first propounded by Maupertius in the eighteenth century, and has been subsequently conceptually sharpened by others.

to one single invariable superlative, such as that of least effort. That, after all, must first be established empirically, as we does with the principle of least action. We can even now note how bizarre the effect would be if a person’s life as one economy according to one dynamic minimum, and at the next moment according to an entirely different dynamic minimum. What would be the effect be for any person if one person’s life were governed throughout by one superlative while his neighbor’s life followed a totally different superlative?

In order to emphasize the ludicrousness of a variety of different superlatives, let us assume that each person consists of two parts, and that each part has a different dynamic superlative of its own. For example, let us assume that one part of the person is governed by least work while the other is governed by least time. In that case the person will represent two distinct problems in dynamics, with the result that he will be, effectively, two entirely different individuals with two distinct sets of dynamical principles. One part of him, in its eagerness to save work, might conceivably spend a long time for the sort of path that, in its eagerness to save time, would minimize one factor and move crosswise without any single dynamic principle behind the total dynamics. For if the person’s entire metabolistic and procreational system is organized so, for the purpose of minimizing work in all its action, then there would have to be a staggering alteration of structure and of functions if the person in question were suddenly to minimize time. Since sudden alterations of such proportions are changes, we are perhaps not overhasty in suspecting a correlation that an individual’s entire activity from birth to death is governed throughout by the same single superlative which, in our opinion, is least effort.

Nor is that is not all. If we remember the extent to which offspring inherit the forms and functions of their parents, we may suspect that this inheritance is possible only if the offspring also inherit the parental dynamic drive that governs the parental forms and functions that are inherited.

Furthermore, if we view the procession of all living forms as the result of slow evolutionary changes from an initial similarity of living matter, then we can understand a fortiori how the one initial single common dynamic superlative might well remain unchanged from generation to generation, regardless of how enormous the changes in forms and functions might become; and that, in turn, will mean that all individuals, regardless of their differences in form and function, will still be governed by the same single superlative.

But though we may argue at length as to the plausibility of one single superlative for all living process, yet when the time comes we shall still need to discover what, in fact, the particular superlative in question is. The actual disclosure of the single hypothetical superlative in question may be difficult for quite obvious reasons. If we take our previous example of the two cities with an intervening mountain chain, in which the paths of least distance, least time, and least work are three different paths, we are obliged in all candor to admit that sometimes one of these paths is taken

and sometimes another. For that matter, a tourist may run up and down the base of the mountain to save distance, while airplanes are flown over the same mountain to save time, while pack horses continue to take a longer and easier winding route. Or, to take another example from daily life, a pedestrian will dart through traffic at considerable risk in order to save time in crossing a street; and sometimes he will take the longer and safer path to the corner, where he will wait for the traffic light. How if we can assert that we are all governed by the same one single dynamic superlative, which apparently varies from occasion to occasion?

But although the superlatives in the foregoing examples seem to be different, are they nevertheless irreconcilable? Before answering this question, let us remember the physicists’ claim that according to their law of falling bodies, all free-standing bodies fall (by least action) to the earth. Yet, despite this claim, we have all observed dry leaves sometimes in the air, or we have seen flocks of the real geese go overhead as if by no means at all they were exceptions to the law of falling bodies. Of course we know from a more careful inspection of the problem that birds and dry leaves are by no means exceptions to the law of falling bodies; on the contrary, if all the factors in the problem are taken into consideration, they are behaving in complete conformity with the law of falling bodies.

May not the same be true of the three different paths to the other side of the mountain? Even though each of these paths may be taken sometimes and by someone, and even though a given person may now take one path and now another, there remains the possibility that the adoption of one course or another by an individual under varying sets of circumstances is governed by the operation of some further single dynamic superlative which forever remains invariant. In any event, we shall argue that such is the case.

More specifically, we shall argue that if we view the above types of situations against the broad background of the individual’s present and future problems, we shall find that an extraordinary expenditure of work at one moment, as an investment in the future as it were, may actually be a temporary devices for reducing the probable rate of the individual’s work expenditure over subsequent periods of his life.

In short, we shall argue that the invariable minimum that governs all varying conduct of all individuals of all living species is the rate of least effort.

III. The Principle of Least Effort

Perhaps the easiest way to comprehend the meaning and implications of the Principle of Least Effort is to show the inadequacy of sheer least work, to which least effort is closely related. This is all the more worth doing because some persons (see below) apparently believe that least work is the basic minimum of living process, as often seems to be the case in particular situations that are considered out of context.

If we remember, however, that an individual’s life continues over a longer or shorter length of time, then we can readily understand how the least work solution to one of his problems may lead to results that will

inevitably increase the amount of work that he must expend in solving his subsequent problems. In other words, the minimizing of work in solving today’s problems may lead to results that will increase tomorrow’s work beyond what would have been necessary if today’s work had been somewhat differently minimized. Conversely, by expending more work than necessary today, one may lessen tomorrow’s probable work.

Since we have argued about the functional relationships of today and tomorrow, we may argue about the functional relationships of the entire succession of events throughout the individual’s whole life, in the sense that all his expenditure of work at one moment may affect the minimizing of his work at a subsequent moment.

In view of the implications of the above provisional considerations, we feel justified in taking the stand that it is the person’s average rate of work-expenditure over time that is minimized in his behavior, and not just his work-expenditure at any one moment in any one isolated problem, with out any reference to its future problems.

Yet a sheer average rate of work-expenditure over time is an entirely meaningful concept, since no mortal can know for certain what his future problems are going to be. The most that any individual can do is to estimate what his future problems are likely to be, and then to govern his present conduct accordingly. In other words, before an individual can minimize his average rate of work-expenditure over time, he must first estimate the probable succession of his future, and then selects a path of least average rate of work through it.

Yet in so doing the individual is no longer minimizing an average rate of work, but a probable average rate of work; or he is governed by the principle of the least average rate of probable work.

For convenience, we shall use the term least effort to describe the preceding least average rate of probable work. We shall argue that an individual’s entire behavior is subject to the minimizing of effort. Or, otherwise said, every individual’s entire behavior is governed by the Principle of Least Effort.

Now that we have described what the Principle of Least Effort is, let us illustrate its operation.

At the risk of being tedious, let the first example be our previous case of the two towns, A and B, that are separated by an intervening mountain range. Here we can see the enormous amount of work that could be saved in travel and trade if the two towns were connected by a tunnel of least distance through the base of the mountain; we can also see the enormous amount of work that it would take to construct such a tunnel. We are simply asked to note that the probable cost in work of digging the tunnel is estimated to be less than the probable work of not having the tunnel; then, if the necessary work force for construction is available, the tunnel will be dug. The problem relates, therefore, to the probable amounts of work involved, as

estimated by one or more persons. Naturally, these persons can have been mistaken in their estimates, with the result that the tunnel can either succeed beyond their wildest hopes, or dismally fail. For we do not deny that “a person’s hindsight is generally better than his foresight.” We merely claim that a person acts on the basis of his “foresight”—with all that that will later be found to imply—according to the Principle of Least Effort.

The above type of argument will also apply to a path of least time over the mountain. Thus the enormous cost of driving highways over the mountain to save time in supplying an army in combat on the other side may be more than justified by the future probable work that is thereby saved.

These cases of the different paths to the other side of the mountain represent decisions of collective action and of collective estimation, since, for example, a tunnel through a mountain is obviously not constructed by a single person but by the collective effort of a great many persons.

And yet we are not restricted to examples of collective effort in illustrating our Principle of Least Effort, which we contend also applies to an individual’s own behavior. We might take the case of a student whose particular path of least effort out of his classroom would send him straight out of the path from his seat to the nearest window, and thence out of the door, through the hall, to the nearest stairway. On the other hand, in the event of a fire, the student might conceivably prefer to rush with least time to the nearest window and adopt a path that is simultaneously a path of least work and of least time and of least distance to the ground! This particular path will be a path of least effort, as estimated by himself, even at the risk of breaking his bones as he leaps—at least, if the alternative is being caught in the path through the smoke-filled corridors. These paths are also paths of least effort, as estimated by the attentive fire-prevention authorities. After all, even the students forgetfully, they can decide which of them, in the light of subsequent events, actually were the shrewdest gamblers in the sense of having both most correctly comprehended the nature and estimated the probabilities of the problem in their lives that was caused by the unexpected fire.

From this second example we can see that the operation of the Principle of Least Effort is contingent upon the mentation of the individual, which in turn includes the operations of “comprehending” the “nature” of elements of a problem, of “assessing their probabilities,” and of “solving the problem.” Since these operations of mentation will claim our consideration at length right here and now, so that we may prepare ourselves for the task of defining mentation, and of showing that the structure and the operations of mentation are also governed throughout by the Principle of Least Effort, since an individual’s mentation is clearly a part of his total behavior, and hence subject to our Principle of Least Effort.

The foregoing examples suffice to illustrate what the Principle of Least Effort is, and what its implications may be for everyday problems. By and large, our explanation of the above commonplace examples was pretty much in line with the way the reader himself would have explained them. We mention this consideration in order to suggest that our chief task may not be that of presenting something that is essentially new so much as that of offering the formal description of a fairly widely felt Gestalt of life as something approaching a basic principle of dynamics.

To avoid a possible verbal confusion, let us note that we are not discussing least probable average rate of work, but a probably least average rate of work.

Chapter Two: On the Economy of Words

As we turn now for the remainder of our study to a demonstration of the Principle of Least Effort, we should keep in mind certain general considerations that will be helpful in guiding our steps. For example we should remember that if Least Effort is indeed fundamental in all human action, we may expect to find it in operation in any human action we might choose to study. In short, any human action will be a manifestation of the Principle of Least Effort in operation, if this Principle is true; therefore all human action is potentially grist for our mill.

In the interest of economy we shall select for our own demonstration first those particular kinds of human action which will most readily admit of the disclosure of the underlying Principle. That is, we shall strive constantly to approach and study our hypothetical Principle from what seems to us to be its most accessible side. For a scientific demonstration can be likened to mountain-climbing—a task in which the mountaineer may either select a path of easiest ascent if he is eager to reach the top, or where he may choose a path of pronounced obstacles if he desires primarily to impress others with his skill. In this study we shall select what seems to be the path of easiest ascent.

Our path is the one that begins with a study of human speech as a set of tools. More specifically, it begins with a study of a vocabulary of words as a set of tools.[1] The reason for selecting this as a beginning is, as we shall see, that the study of words offers a key to an understanding of the entire speech process, while the study of the entire speech process offers a key to an understanding of the personality and of the entire field of biosocial dynamics. Hence the contents of the present chapter will be of crucial importance for our entire study because in this chapter we shall untie a knot that we shall find duplicated again and again in other biosocial phenomena. The care and completeness with which we untie this first knot will render all future knots so much the easier to untie.1

I. In Medias Res: Vocabulary Usage, and the Forces of Unification and Diversification

Man talks in order to get something. Hence man’s speech may be likened to a set of tools that are engaged in achieving objectives. True, we do not yet know that whenever man talks, his speech is invariably directed to the

Human speech is traditionally viewed as a succession of words to which “meanings” (or “usages”) are attached. We have no quarrel with this traditional view which, in fact, we here adopt. Nevertheless in adopting this view of “words with meanings” we might profitably combine it with our previous view of speech as a set of tools, and state: words are tools that are used to convey meanings in order to achieve objectives.

Yet once we say that words are tools, we broach thereby the question of the possible economies of speech; and as soon as we inquire into the possible economies of speech we remember that the sheer ability to speak at all represents an enormous convenience in present-day human social activity, whereas the inability to speak is a signal handicap. Since both the conveniences of being able to speak, and the handicap of being unable to do so, refer admittedly to the saving of effort, we may say that there is a potential general economy in the sheer existence of speech, in the sense that some human objectives are more easily obtained with speech than without it. The case is similar to that of a set of carpenter'''s tools whose sheer existence may be said to have a potential general economy for the carpenter.

But beyond this potential general economy of speech there are further possibilities for economy in the manner in which speech is used. For if speech consists of words that are tools which convey meanings, there is the possibility both of a more economical way, and of a less economical way, to use word-tools for the purpose of conveying meanings. Hence in addition to the general economy of speech there exists also the possibility of an internal economy of speech.

Now if we concentrate our attention upon the possible internal economies of speech, we may hope to catch a glimpse of their inherent nature. Since it is usually felt that words are “combined with meanings” we may suspect that there is latent in speech both a more and a less economical way of “combining words with meanings,” both from the viewpoint of the speaker and from that of the auditor.2

From the viewpoint of the speaker (the speaker'''s economy) who has the job of selecting not only the meanings to be conveyed but also the words that will convey them, there would doubtless exist an important latent economy in a vocabulary that consisted exclusively of one single word—a single word that would mean whatever the speaker wanted it to mean. Thus if there were different meanings to be verbalized, this word would have different meanings. For by having a single-word vocabulary the speaker would be spared the effort that is necessary to acquire and maintain a large vocabulary and to select particular words with particular meanings from this vocabulary. The single-word vocabulary, which reflects the speaker'''s economy, may be likened to an imaginary carpentry kit that consists of a single tool of such art that it can be used exclusively for all the different tasks of sawing, hammering, drilling, and the like, thereby saving the labor of otherwise devising, maintaining, and using a more elaborate toolage.

But from the viewpoint of the auditor (the auditor'''s economy), a single-word vocabulary would represent the acme of verbal labor, since he would be faced by the impossible task of determining the particular meaning to which the single word in a given situation might refer. Indeed from the viewpoint of the auditor, who has the job of deciphering the speaker'''s meanings, the important internal economy of speech would be found rather in a vocabulary of such size that it possessed a distinctly different word for each different meaning to be verbalized. Thus if there were different meanings, there would be different words, with one meaning per word. This one-to-one correspondence between different words and different meanings, which represents the auditor'''s economy, would save effort for the auditor in his attempt to determine the particular meaning to which a given spoken word referred.3

As far as the problem of words and meanings is concerned, we note the presence of two far-reaching contradictory economies that relate in each case to the number of different meanings that a word may have. Thus if there are an number of different distinctive meanings to be verbalized, there will be (1) a speaker'''s economy in possessing a vocabulary of one word which will refer to all the distinctive meanings; and there will also be (2) an opposing auditor'''s economy in possessing a vocabulary of different words with one distinctive meaning for each word. Obviously the two opposing economies are in extreme conflict.

We may even visualize a given stream of speech as being subject to two “opposing forces.” The one “force” (the speaker'''s economy) will tend to reduce the size of the vocabulary to a single word by unifying all meanings behind a single word; for that reason we may appropriately call it the Force of Unification. Opposed to this Force of Unification is a second “force” (the auditor'''s economy) that will tend to increase the size of a vocabulary to a point where there will be a distinctly different word for each different meaning. Since this second “force” will tend to increase the diversity of a vocabulary, we shall henceforth call it the Force of Diversification. In the language of these two terms we may say that the vocabulary of a given stream of speech is constantly subject to the opposing Forces of Unification and Diversification which will determine both the number of different words in the vocabulary, and also the meanings of those words.

In adopting the term force to describe the two opposite economies that

are hypothetically latent in speech, we must remember that the term refers to what people will in fact do and not to what they are at liberty to do if they wish. For we are arguing that people do in fact always act with a maximum economy of effort, and that therefore in the process of speaking-listening they will automatically minimize the expenditure of effort. Our Forces of Unification and Diversification merely describe two opposite courses of action which from one point of view or the other are alike economical and permissible and which therefore from the combined viewpoints will alike be adopted in compromise. From this it follows that whenever a person uses words to convey meanings he will automatically try to get his ideas across most efficiently by seeking a balance between the economy of a small wieldy vocabulary of more general reference on the one hand, and the economy of a larger one of more precise reference on the other, with the result that the vocabulary of different words in his resulting flow of speech will represent a vocabulary balance between our theoretical Forces of Unification and Diversification.4

II. The Question of Vocabulary Balance

We obviously do not yet know that there is in fact such a thing as vocabulary balance between our hypothetical Forces of Unification and Diversification, since we do not yet know that man invariably economizes with the expenditure of his effort; for that, after all, is what we are trying to prove. Nevertheless—and we shall enumerate for the sake of clarity—if (1) we assume explicitly that man does invariably economize with his effort, and if (2) the logic of our preceding analysis of a vocabulary balance between the two Forces is sound, then (3) we can test the validity of our explicit assumption of an economy of effort by appealing directly to the objective facts of some samples of actual speech that have served satisfactorily in communication. Insofar as (4) we may find therein evidence of a vocabulary balance of some sort in respect of our two Forces, then (5) we shall find ipso facto a confirmation of our assumption of (1) an economy of effort. Therefore much depends upon our ability to disclose some demonstrable cases of vocabulary balance in some actual samples of speech that have served satisfactorily in communication.

Fortunately, if a condition of vocabulary balance does exist in a given sample of speech, we shall have little difficulty in detecting it because of the very nature and direction of the two Forces involved. On the one hand, the Force of Unification will act in the direction of decreasing the number of different words to 1, while increasing the frequency of that 1 word to 100%. Conversely, the Force of Diversification will act in the opposite direction of increasing the number of different words, while decreasing their average frequency of occurrence towards 1. Therefore number and frequency will be the parameters of vocabulary balance.

Since the number of different words in a sample of speech together with their respective frequencies of occurrences can be determined empirically, it is clear that our next step is to seek relevant empiric information about the number and frequency of occurrences of words in some actual samples of speech.

A. Empiric Evidence of Vocabulary Balance

James Joyce’s novel Ulysses, with its 260,430 running words, represents a sizable sample of running speech that may fairly be said to have served successfully in the communication of ideas. An index to the number of different words therein, together with the actual frequencies of their respective occurrences, has already been made with exemplary methods by Dr. Miles L. Hanley and associates who have quite properly argued that all words are different which differ in any way “phonetically” in the fully inflected form in which they occur (thus the forms, give, gives, gave, given, giving, giver, gift represent seven different words and not one word in seven different forms).5

To the above published index has been added an appendix from the careful hands of Dr. M. Joos, in which is set forth all the quantitative information that is necessary for our present purposes. For Dr. Joos not only tells us that there are 29,899 different words in the 260,430 running words; he also ranks those words in the decreasing order of their frequency of occurrence and tells us the actual frequency, , with which the different ranks, , occur. By consulting this appendix we find, for example, that the 10th most frequent word () occurs 2,653 times ()f equals 2,653; or that the 100th word ()r equals 100 occurs 265 times ()f equals 265. In fact, the appendix tells us the actual frequency of occurrence, f, of any rank, r, from r equals one to r equals 29,899, which is the terminal rank of the list, since the Ulysses contains only that number of different words.

It is evident that the relationship between the various ranks, r, of these words and their respective frequencies, f, is potentially quite instructive about the entire matter of vocabulary balance, not only because it involves the frequencies with which the different words occur but also because the terminal rank of the list tells us the number of different words in the sample. And we remember that both the frequencies of occurrence and the number of different words will be important factors in the counterbalancing of the Forces of Unification and Diversification in the hypothetical vocabulary balance of any sample of speech.

Turning to the quantitative data of the Hanley Index we can see from the arbitrarily selected ranks and frequencies in the adjoining Table 2-1 that the relationship between r and f in Joyce’s Ulysses is by no means haphazard. For if we multiply each rank, r, in Column I of Table 2-1 by its corresponding frequency, f, in Column II, we obtain a product, C, in Column III, which is approximately the same size for all the different ranks and which, as we see in Column IV, represents approximately one tenth of the

23

260,430 running words which constitute the total length of James Joyce’s Ulysses. Indeed, as far as Table 2-1 is concerned, we have found a clearcut correlation between the number of different words in the Ulysses and the frequency of their usage, in the sense that they approximate the simple equation of an equilateral hyperbola:

in which r refers to the word’s rank in the Ulysses and f to its frequency of occurrence (as we ignore for the present the size of C).

TABLE 2-1 — Arbitrary Ranks with Frequencies in James Joyce’s Ulysses (Hanley Index)

I — Rank ()II — Frequency ()III — Product of I and II ()IV — Theoretical Length of Ulysses ()
102,65326,530265,500
201,31126,220262,200
3092627,780277,800
4071728,680286,800
5055627,800278,800
10026526,500265,000
20013326,600266,000
3008425,200252,000
4006224,800248,000
5005025,000250,000
1,0002626,000260,000
2,0001224,000240,000
3,000824,000240,000
4,000624,000240,000
5,000525,000250,000
10,000220,000200,000
20,000120,000200,000
29,899129,899298,990

The data of this table give clear evidence of the existence of a vocabulary balance.

We must not forget that Table 2-1 contains only a few selected items out of a possible 29,899; hence the question is legitimate as to the possible rank-frequency relationship between the rest of the 29,899 different words. Although we cannot easily present in tabular form the rank-frequency relationships of all these different words, we nevertheless can present them quite conveniently on a graph, because we know that the equation, r times f equals C, will appear on doubly logarithmic chart paper as a succession of points descending in a straight line from left to right at an angle of 45°. And if we plot the ranks and frequencies of the 29,899 different words on doubly

24

logarithmic chart paper, and if the points fall on a straight line descending from left to right at an angle of 45°, we may argue that the rank-frequency distribution of the entire vocabulary of the Ulysses follows the equation r times f equals C, and suggests the presence of a vocabulary balance throughout.

As to the details of the graphical plotting of this particular equation (which will be repeated again and again throughout our study) we shall plot successive ranks from 1 through 29,899 horizontally on the -axisx axis, or abscissa. Then, in measuring frequency on the -axisy axis, or ordinate, we

Log-log plot of rank versus frequency with three descending curves at approximately 45°: curves labeled A, B, and CFig. 2-1. The rank-frequency distribution of words. (A) The James Joyce data; (B) the Eldridge data; (C) ideal curve with slope of negative unity.

shall give for each rank a dot which corresponds to the actual frequency of occurrence of the word of that rank. After we have completed the plotting of the actual frequencies of our 29,899 ranked words, we shall connect the dots with a continuous line in order to note whether the line is straight and whether it descends from left to right at the expected angle of 45°.

In Fig. 2-1 we present in Curve A the data of the entire Ulysses thus plotted, and the reader can assess for himself the closeness with which this curve descends from left to right in a straight line at an angle of 45°. In order to suggest that the Ulysses is not unique in respect of a hyperbolic rank-frequency word distribution, we include gratuitously in Curve B of Fig. 2-1 the rank-frequency distribution of the 6,002 different words in fully inflected form as they appear in a total of 43,989 running words of

25
combined samples from American newspapers as analyzed by R. C. Eldridge.[^6] Curve C is an ideal curve of 45° slope that has been added to aid the reader’s eye.

We note that the curves of Fig. 2-1 conform with considerable closeness to a straight line with the expected slope of 45°, except for the emergence of “steps” of progressively increasing size as the line approaches the bottom. Although we shall shortly see that these “steps” result from integral frequencies and are governed by the equation, r times f equals C, we may now only say that the data confirm our equation merely down to where the steps begin. However, we note that an extension of the straight line through the “steps” would in most cases cut them fairly squarely through the middle (for reasons to be explained later), and that therefore the “steps” are by no means capricious in occurrence but have an orderliness of their own that is clearly not unrelated to the orderliness of the straight line above.

B. The Significance of

Before discussing the reasons for the emergence of the “steps” in Fig. 2-1, let us dwell briefly upon the significance of the curves themselves which clearly show that the selection and usage of words is a matter of fundamental regularity of some sort of an underlying governing principle that is not inconsistent with our theoretical expectations of a vocabulary balance as a result of the Forces of Unification and Diversification.

Perhaps the easiest way to appreciate the fundamental regularity exhibited by our curves is to ignore for the moment how they do appear and to inquire instead how they might appear if no underlying governing principle were involved. In short, let us inquire into the various ways that a rank-frequency distribution both could, and could not, appear from the particular manner in which we are plotting the data so that we may see how remote the probabilities are of their conforming to the rectilinear distribution we have observed.

In the first place, since we are ranking the words from left to right in the decreasing order frequency, it is evident that the line that connects the succession of dots can at no point bend upwards, since an upward bend at any point would indicate an incorrect ranking of the data according to decreasing frequencies. On the other hand, the line can and, in fact, will proceed horizontally whenever adjacent ranks have precisely the same frequencies (as happens to be the case with the horizontal lines of the “steps” at the bottom of the curves of Fig. 2-1, as we shall presently see). Hence we may predict in advance that any rank-frequency distribution may never slope upward from left to right although it may be horizontal. But that is not all. We may also predict that a rank-frequency curve will never bend downwards in a true vertical, since the line must pass from left to right in order to connect the dots of adjacent ranks. The apparently vertical lines of the “steps” of Fig. 2-1 are not truly vertical, since they do in fact connect adjacent dots. On the other hand, as long as the line never becomes a true vertical, it can bend downwards with any slope at any point.

As far as our method of plotting our data is concerned, we may say in

26

advance that the line proceeding from left to right in a rank-frequency distribution may twist and turn at any point on the graph paper as long as it never bends upwards and never bends downwards in a true vertical. In this connection the reader might take a pencil and paper and draw lines of various configurations and contortions that connect the upper left-hand corner with the lower right-hand corner—lines that avoid upward bends and true verticals—in order to assure himself of the vast number of possibilities that lie within the restrictions of our method of plotting. After completing his “random lines” the reader will appreciate the orderliness of the lines of Fig. 2-1; and he will see how this orderliness points to the existence of a fundamental governing principle that determines the number and frequency of usage of the words in the stream of speech, regardless of whether or not the speakers and auditors are aware of the existence of the principle, and regardless of whether or not our Forces of Unification and Diversification in vocabulary balance provide a necessary explanation of it. Since all the words of Fig. 2-1 had “meanings” in their respective samples, the reader may infer from the orderliness of the distribution of words that there may well be a corresponding orderliness in the distribution of meanings because, in general, speakers utter words in order to convey meanings.

III. The Orderly Distribution of Meanings

Taking a temporary leave of the distribution of words in Fig. 2-1, let us now turn our attention to the question of the distribution of the meanings of words. We have previously argued that under the conflicting Forces of Unification and Diversification the m number of different meanings to be verbalized will be distributed in such a way that on the one hand no single word will have all m different meanings and that on the other hand there will be fewer than m different words. As a consequence, we may expect that at least some words must have multiple meanings. There remains then the problem of determining, first, which words will have multiple meanings and, second, how many different meanings these words of multiple meaning will have. In the solution of this problem, the Forces of Unification and Diversification will stand us in good stead.

Let us begin by turning our attention to the most frequently used word in the stream of speech, with special reference to the actual samples of Fig. 2-1. We shall arbitrarily designate the frequency of this most frequent word with the letter, F sub one. The question now remains as to the m sub one number of different meanings which are represented by F sub one. And here we may say that, regardless of the size of m sub one, if we multiply m sub one by f sub one, which represents the average frequency of occurrence of the m sub one meanings, we shall obtain F sub one, since F sub one is made up of the total frequencies of its different meanings. Therefore we may write:

With this simple equation in mind, let us recall our previously discussed Forces of Unification and Diversification and inquire into their respective

27
influences upon the sizes of m sub one and f sub one. Obviously, the Force of Unification which theoretically acts in the direction of putting all different meanings behind a single word will tend to increase the size of m sub one at the expense of the size of f sub one. On the other hand, the Force of Diversification which theoretically acts in the direction of reducing the number of different meanings per word will tend to increase f sub one at the expense of m sub one. Therefore the respective sizes of m sub one and f sub one of our previous equation will again represent the action of the opposing Forces of Unification and Diversification.

Of course, we do not know a priori what the comparative strength of these two Forces may be. Yet we have observed from the data of Fig. 2-1 that there is a hyperbolic relationship between the n number of different words in the samples and their respective frequencies of occurrence. Therefore we may suspect that our two Forces of Unification and Diversification stand, in general, in a hyperbolic relationship to one another, with the result that m sub one and f sub one will also stand in a hyperbolic relationship with one another, with the further result that m sub one will tend to equal f sub one.

However if m sub one equals f sub one and since m sub one times f sub one equals F sub one, then clearly m sub one will equal the square root of F sub one, or square root of F sub one.

But now let us note that the above argument will apply mutatis mutandis to the m sub r number of different meanings of the word whose comparative frequency of occurrence is F sub r, with the result that the following simple equation may be expected:

This simple equation is of interest, for it means that if (1) we make a rank-frequency distribution of the words of a sample of speech, as was done for the Ulysses and Eldridge data of Fig. 2-1, and if (2) we find that this distribution yields the straight line of an equilateral hyperbola as found in Fig. 2-1, then (3) we may conclude from the nature of the above argument and equation that a rank-frequency distribution of the different meanings of those words on doubly logarithmic paper would yield a straight line descending from left to right to the point, X equals n, yet intercepting only one half as much on the -axisy axis as on the -axisx axis (that is, it will have what is technically called a negative slope of one half, or of .5). The reason for this is that the m sub r number of different meanings for each of the -rankedr ranked words will be represented on doubly logarithmic paper by a point that is in each case one half of the F sub r of the respective ranked words. We shall call this the theoretical law-of-meaning distribution.

To determine empirically whether this theoretical law-of-meaning distribution exists, we could take the data of Fig. 2-1 and, after consulting a suitable dictionary, we could graph the m sub r number of different meanings for each r different word, and note the resulting meaning-frequency distribution. The resulting meaning-frequency distribution would refer only to the particular Ulysses and Eldridge word-frequency distributions, and therefore would lack a more general applicability.

It would be of more general applicability and equally valid for our purposes if we selected the more comprehensive word-frequency distribution of

28

English as made and published by E. L. Thorndike on the basis of a count of 10 million running words.[^7] Although Dr. Thorndike has published only the 20,000 most frequent words of his count, nevertheless these 20,000 words will represent the average frequencies of standard English better than the particularized vocabularies of the data of Fig. 2-1. It is true that Dr. Thorndike has for the most part ignored the inflectional endings of words; instead he has subsumed the frequencies of occurrence of practically all different inflectional forms of a given word under the dictionary form of that word (i.e., he used what is technically known as a lexical unit); however we have no reason to suppose that any “law of meanings” would be seriously distorted if we concentrated our attention upon lexical units and simply ignored variations in number, case, or tense. Nor need we be disturbed by the fact that Dr. Thorndike did not list the actual frequencies of the different words but merely noted the 1st thousand most frequent, the 2nd thousand most frequent, and so on down through the 20th thousand most frequent, with a further notation of whether a given word of the first 5000 words was among the first or second 500 words of its respective thousand. This lack of a precise numerical notation—far from invalidating his count—offers a genuine challenge to our thesis. For (1) if we are correct in generalizing upon the data of Fig. 2-2 by stating that the distributions are representative of English, and (2) if our theoretical law-of-meaning distribution be correct, then we may suspect, both (3) that Thorndike’s 20,000 words would follow a hyperbolic rank-frequency distribution of words and (4) that the distribution of meanings of the 20,000 words when plotted on doubly logarithmic graph paper will yield a negative slope of .5 as previously explained. Therefore we may test our theoretical law-of-meaning distribution by turning directly to an analysis of the average m number of different meanings per word in each of the 20 successive sets of one thousand words.

Fortunately for the analysis of the meanings of the 20,000 words, we have available the Thorndike-Century Dictionary which selected the m different meanings to be presented for each word (except for the 500 most frequent) on the objective basis of Dr. Irving Lorge’s The English Semantic Count.[^8] Hence the m number of actually used different meanings for each word in the dictionary has been determined empirically, with the result that in making our meaning-frequency analysis we need not fear including archaic or obsolescent meanings which might well distort our distribution.

Thanks to the help of some of my students, who undertook the task of noting the number of different meanings in Thorndike’s dictionary for each of the 20,000 words of the list, we present in Fig. 2-2 the average number of meanings per word (on the ordinate) for each successive set of 1000 words on the abscissa. Since the average number of meanings per word in each thousand refers in fact to the 500th word (or class-middle) of each thousand, the points on the abscissa represent these class-middles in all cases; that is, they represent the values of the 500th, 1500th, 2500th, … 19,500th words respectively.

A glance at the data of Fig. 2-2 suffices to show that the points descend

29
in a strikingly straight line which is not far off from our theoretically expected negative slope of .5 (i.e., negative point five). If we calculate by least squares the slope of the best straight line through the points, we arrive at the value negative point four six zero five ()plus or minus point zero zero eight three with the -intercepty intercept at (antilog). This calculated value is not far off from our expected negative point five slope.6

The approximation may be even closer than that if we remember that The English Semantic Count was not used for the 500 most frequent words (whose differentiation of meanings is truly difficult for reasons that will be apparent in our following chapter). Because of this consideration the first point at the left of our chart is suspect. If we ignore it and recalculate the slope for the remaining 19 points we have a slope of negative point four six five six ()plus or minus point zero zero two seven, which is slightly nearer to the expected negative point five slope.

Log-log plot of meanings (average) versus rank (in thousands), showing points descending in a straight line at approximately negative half slopeFig. 2-2. The meaning-frequency distribution of words.

If we turn our attention now to the 10 successive sets of 500 words which constitute the 5000 most frequent words in the list, and if we again ignore the suspected first 500 words for reasons already presented, we have a slope of negative point four eight nine nine ()plus or minus point zero zero three, with which we may scarcely quarrel as an approximation of a negative point five slope.[^10]

It is of course regrettable that additional sets of data on this important point are not available. Nevertheless the results of even this one study are so striking that pending the future findings of empiric analysis we are not rash in concluding that a law-of-meaning distribution exists according to which the m average number of meanings per word of a thousand words (when ranked in the order of decreasing frequency) will equal the square

30

root of the average frequency of the words’ occurrence (or will decrease according to the square root of the rank).

Although later we shall again return to the entire question of the “meanings of words” with the problem of defining the term meaning,[^11] we may even now feel that our theoretical Forces of Unification and Diversification have led us to the empiric disclosure not only of a simple equation for the distribution of words (in the form, r times f equals C, with r an integer) but also of a simple equation for the distribution of the meanings of those words which may be put down in the form of the equation, m sub r equals the square root of F sub r, in which F sub r equals C sub r.

Incidentally, the fact that we have no actual rank-frequency distribution for the 20,000 words of Dr. Thorndike’s frequency list does not invalidate our above conclusion; on the contrary, we shall present so many word-rank-frequency distributions in our following pages that the reader will be more than ready to believe that, if we had a rank-frequency distribution of the 20,000 most frequent words of the Thorndike analysis, it would probably be rectilinear like those of Fig. 2-1, at least for the first 10 or 12 thousand most frequent words.

With the assurance of the law-of-meaning distribution of Fig. 2-2, let us return now to a study of the significance of the rectilinear distributions of Fig. 2-1, which descend with a negative slope of 1 except for the “steps” of increasing magnitude at the bottom, and which indicate the existence of a vocabulary balance.

31

Footnotes

  1. For the sake of simplification we shall use the term least effort in the present chapter to apply not only to situations of least probable work, but also to situations in which the argument is restricted to immediate behavior, which is technically one of least work.
    attainment of objectives. Nevertheless it is thus directed sufficiently often to justify our viewing speech as a likely example of a set of tools, which we shall assume to be the case.

  2. Nor does the word need to be spoken; it may also be written. The situation of the writer-reader is analogous to that of the speaker-auditor in respect of internal economies of usage of words, even though a reader is not so immediately present to a writer as an auditor is to a speaker, and even though the word-usage of written speech may differ somewhat from that of spoken speech for reasons that we shall scout in a later chapter. If we continue for the time being to discuss words without dichotomizing between written and spoken verbalizations, we do so in the interest of a legitimate simplification which seems to be justified at the beginning of our analysis of words and their usage as we think the reader will agree upon reflection.

  3. Later we shall define a meaning of a word as a kind of response that is invoked by the word.

  4. We shall consistently capitalize the terms, Force of Unification and Force of Diversification, in order to remind ourselves that these Forces do not represent forces as physicists traditionally understand the term, but only the natural consequences of our assumed underlying economy of effort. Moreover our term balance will include what are technically known as steady states and the equilibria of the physicist and of the economist.

  5. [footnote reference 2 in source]

  6. This slope is probably the most reliable, since it refers to the most frequent 10,000 words that are likely to be found in an optimum sample of 100,000 running words. For a discussion of an optimum sample see below.