Javascript

Showing posts with label mathematics. Show all posts
Showing posts with label mathematics. Show all posts

Monday, July 26, 2010

media hype and bell's theorem

So, I found the following e-mail dated 2007-08-24, in which I replied to an e-mail forward titled "FW: We Have Broken The Speed Of Light". The forward was regarding physicists from the University of Koblenz demonstrating quantum-entanglement, but the reporter chose to refer to that as "We Have Broken The Speed Of Light".


It's just a misleading headline, and the German scientists have done nothing new. The news agencies are renowned for pulling this sort of stuff -- ignoring scientific progress which inches along like a slug on the pavement, and then after decades of ignoring this, some news agency realizes that the slug has moved far enough from where the public was last informed it was, and the agency now declares this a "leap" in scientific understanding with a sensationally misleading headline and complete misunderstanding of the fundamentals. Unfortunately, most people do not grasp quantum mechanics and so any reporter can rise to fame with whatever sort of babble he writes.

The public would be more skeptical if, for example, the reporter claimed that the new Toyota manages to draw their new cars with 2,000 actual horses. That's not what a horsepower is, and people know that. We know that horsepower is a unit of force approximately equal to the output of one horse. This is important because the output of one horse is not the same thing as a horse, because one is a unit of energy per distance (force) that can be produced by a reasonably small device, and the other is an equestrian creature that needs to be fed and when in quantities exceeding twenty will occupy all lanes of a typical highway and would struggle to fit inside the average consumer's garage.

Physics isn't the only realm where "new"s is really "old"s repackaged under the mantra "if you haven't seen it, it's new to you." Remember 9/10/2001? The summer was all about shark attacks, and every news agency jumped on that band wagon because shark attacks went un-reported so long that it became new and terrifying. Of course, the little known fact was that 2001 witnessed a below average incidence of annual deaths due to sharks. That's right, a year with fewer shark attacks than usual fueled a media frenzy over each and every shark attack.

Take, for example, these headlines:

Why Can't We Be Friends? A horrific attack raises old fears, but new research is revealing surprising keys to shark behavior
(TIME magazine, cover, 2001)
A Scary Jump in Shark Attacks...Could Threaten the Sharks (Businessweek, April 23, 2001)
Expert confirms surfer was bitten by shark (CNN, July 20, 2001)
Boy dies after shark attack (CNN, Sept 2, 2001)
Man killed in N.C. shark attack; woman hurt (CNN, Sept 3, 2001)

Any many more..

Contrast that media hype with academia trying to bring people to their senses:

"GAINESVILLE, Florida, February 18, 2002 (ENS) - Despite the prevailing perception that 2001 was a banner year for shark attacks, actual numbers were slightly down, a new University of Florida study shows." -- University of Florida (ufl.edu)

" 'Falling coconuts kill 150 people worldwide each year, 15 times the number of fatalities attributable to sharks,' said George Burgess, Director of the University of Florida's International Shark Attack File and a noted shark researcher." -- Daily University Science News (unisci.com) May 23, 2002

Dubious Data Award 2001: "The frenzy was remarkable. According to the NEXIS database, there were a mere 58 stories about shark attacks in the US print media in June. This increased tenfold in July to 592, and it rose again in August to 684. September was the month the story would have consumed all others, with 624 entries up to and including September 11. The advent of another dangerous but unseen villain stopped all that. ... During the 1990s, when only five people were killed by sharks, 28 children were killed by falling TV sets. The Times editorial mentioned above concluded from our data that, loosely speaking, 'watching Jaws on TV is more dangerous than swimming in the Pacific.' " (stats.org, Jan 1, 2002)

Right, so we've thoroughly exposed how most of these headlines are useless "old"s repacked with bad math and poor editorializing to shock you as "new"s. Why is the faster-than-light article misleading? Well, first, we need to delve into what speed is. Speed is distance per time. What is distance? One might answer "the distance between A and B is the length of a straight line between those points". But that's not quite right. A straight line is the definition of distance only in Cartesian space, but in other geometries where the straight line is sub-optimal (longer than the shortest path) or impossible (shorter than the shortest path), the length of a "straight line" is not the distance. So what is the distance? It's the length of the geodesic -- which is just a math term for the "shortest possible path". Aha! So now we're making progress - that means that I can travel at a speed (distance per time) slower than light but arrive sooner by reducing the distance! But of course we all knew that -- we find "short-cuts" to drive to work, often traveling on slow back-roads instead of fast superhighways yet still arriving quicker.

So how do these spatial short-cuts work? Well, there are all kinds of spatial short-cuts, because physicists have long since known that spacetime is heavily warped with many microfissures like quantum entanglements peppering the universe as well as with large fissures made by gravity wells such as blackholes, neutron stars, and wormholes. Physics of the very big (astrophysics) is only a looking glass, but physics of the very small (quantum mechanics) prescribes us experiments which are feasible to undertake. Specifically, the easiest way we can warp spacetime is through quantum entanglement -- which is far more than just a theory since Bell's paper shook the scientific community in 1964 with results which we will eventually get to in this email.

Okay, but what is quantum mechanics? Back in Einstein's day, it was centered on one notion: that a particle could be indeterminate -- that is, not only is its state unknown by you and me, but its state is unknown by even itself and the universe and, if we want to drag theology into this, God. Einstein famously exclaimed "God does not play dice!" to ridicule the notion of indeterminate states. Who were Einstein's intellectual nemeses? Neils Bohr and Werner Heisenberg, the two who came up with what became known as the Copenhagen interpretation while working together in Copenhagen, 1927. Essentially, Bohr and Heisenberg believed that sub-atomic particles were probabilistic and not deterministic. Einstein believed that probabilities were merely mathematical constructs on paper but actual things in the universe were deterministic.

Unlike a purely philosophical disagreement, this disagreement was scientific. If particles were determinate, they would behave like you would expect. If particles could be indeterminate, then some peculiar things can be done: particle X can be at A with 50% probability and at B with 50% probability -- which is a very different thing from saying X is at A 50% of the time and at B 50% of the time. To illustrate the difference, suppose I am (A) dialing my apartment number with my cell phone with 50% probability and I am (B) dialing my cell phone with my apartment number with 50% probability. There is just one (My Name – edited out), so you would expect either the apartment phone to ring or my cell phone to ring. However, that interpretation (Einstein's) would assume that (A) happens 50% of the time and (B) happens 50% of the time but never both. Under the Copenhagen interpretation, you would sometimes hear the apartment phone ring (A), you would sometimes hear the cell phone ring (B), and you would sometimes hear a busy signal (A+B)! That third possibility is highly important -- it means that both (A) and (B) occurred simultaneously and interacted with itself. How on earth can I get a busy signal by calling one of my phones from the other? You can see why there was a lot of debate - quantum mechanics does not sound reasonable. The scientific community was split.

All this is purely theoretical at the time (1920s and 1930s), but people were intrigued. So what is quantum entanglement? Well, "entanglement" was a term Schrodinger used during his debates with Einstein after the 1935 EPR paper (a paper written by Einstein, Podolsky and Rosen showing bizarre conclusions derived from quantum mechanics, a thinly veiled attempt at demonstrating the absurdity of such a view in a rational universe). Shortly put, these indeterminate states can be determinate by entangling with it. Schrodinger famously used his "cat in a box" thought experiment. We start with a single radioactive atom with a 1 week half-life, which means that it will decay in the first week with equal probability that it will not decay. Second, we add an execution machine which releases poisonous gas, triggered by alpha/beta emission (radioactive decay). Third, we add a living cat. Fourth, we seal them all in one impenetrable box. The cat's fate is entangled with the fate of the radioactive atom. If the atom decays, the cat dies, else the cat lives. However, because of the how the psi wave function works (the mathematical function calculating probabilities for indeterminate states), only the act of observing will collapse the wave function into a single determinate state, but for only the observer and no one else. The cat is observing the execution machine (through being alive or dead), and the execution machine is observing the radioactive atom (through being triggered or not). Therefore, the radioactive atom is in a determinate state for the machine and the cat. However, we the people on the outside are observing none of this, so the radioactive atom is in a quantum superposition of both states. Therefore, the cat also has entered quantum entanglement, because it observed the quantum particle yet was not itself observed. This means that the cat is also in quantum superposition, one involving both living and dead states. If we observe the radioactive atom, we make determinate the state of the cat; similarly, if we observe the cat first, then we make determinate the state of the radioactive atom. This is entanglement.

The Copenhagen interpretation says that such a thing is possible. The classical approach, endorsed by Einstein, says such a thing is impossible -- that it cannot possibly be that only when the box is opened the atom and cat decide to be either decayed and dead or non-decayed and alive. The classical approach suggests that this decision happened, that either it was fated that the atom decay and the cat die or it was fated that the atom not decay and the cat remain alive. The classical approach could never accept that a cat dead for one entire day could be decided now at the time of opening the box to have died yesterday. The present affecting the past is considered counterfactual, and this was the crux of the EPR paradox showing how quantum mechanics is in direct contradiction with a rational universe.

As reasonable as the classical interpretation is and as absurd as the Copenhagen interpretation, and quantum mechanics along with it, sounded, the mathematics was inescapable and particle physicists had known since the Thomas Young double-split experiment in 1801 that physics of the very-small behaves in absurd ways. While the rift within the scientific community lingered, exacerbated by Einstein's obstinacy at pursuing a feud with quantum mechanics, many bright minds and leaders of the community flocked toward quantum mechanics. In many ways, Einstein's opposition was a bit hypocritical. His own theory of special relativity, and later general relativity, battled against classical Newtonian physics, and the new theories of quantum mechanics did not jeopardize relativity -- in fact, quantum mechanics and the theory of relativity can coexist side by side perfectly without contradictions. The dispute was academic, it was abstract, it was about the philosophical nature of reality. It was, that is, until 1965 when Irish physicist John Bell released a paper which thundered through the community shaking classical beliefs down to their very core and proving beyond any doubt that the "rational universe" envisaged by classical notions was what was fake, rendering any classical notions of local realism illusory.

Specifically, Bell's theorem proved how local hidden variables cannot explain the phenomena observed at the quantum scale. It sounds like a benign statement, unless one realizes what local hidden variables are. In the example with Schrodinger's cat in a box, the classical approach would suggest that whether the atom decayed and the cat died was hidden, not indeterminate. Classical thinking states that this special impenetrable box merely hides what happened, but all the events still transpired and are not in an indeterminate state. Bell's theorem shows that however reasonable such classical thinking is, it is wrong, hidden local variables do not work. The theorem is simple, and involves setting up an apparatus to empirically verify how our universe isn't a rational universe, one which could be explained merely by local hidden variables.

The remainder of this e-mail is devoted to summarizing Bell's theorem:

Bell's theorem begins, if local realism is true, then there is no impact of a distant actor's actions when measuring a local particle's spin. The phenomenon where two particles under quantum entanglement can have perfect correlation when measured at identical angles (let X be this angle) is a phenomenon that can be explained under the model of local realism by employing local hidden variables, as follows: A deterministic mapping could be programmed into these entangled particles at the time they were entangled, which is feasible because at the time they were entangled they were proximate to each other. The deterministic mapping, call it m, would map an angle, from 0 to 360 degrees, to spin, -1 for counter-clockwise and +1 for clockwise. Therefore, even if these entangled particles are now apart by great distances, one actor measuring one of the particles at angle X would see a spin m(X) and a distant actor measuring the other particle at angle X would also see spin m(X). Since m is a local property carried by each particle, nothing non-local affects the measurement.

While local realism seems to work at explaining entanglement, Bell's theorem devises a situation where it fails. Let Q and Q' be two angles the distant actor can measure at, and R and R' be two angles the local actor can measure at. Therefore, we need concern ourselves with only a reduced m, one which doesn't handle 0 to 360 degrees but instead only four angles, Q, Q', R, and R'. Since m is deterministic, there are only 16 possible mappings it could be -- e.g. m(Q,Q',R,R') could be (1,1,1,1) or (1,1,1,-1) or (1,-1,-1,1) et cetera. Let M be the set of these 16 possible mappings. Let s be the spin the local actor observes and let t be the spin the remote actor observes. We can decompose E[s*t|Q,R], the expected value for the product s*t for when angles Q and R are chosen but m is free to be anything in M, into "sum p(m)*m(Q)*m(R) for all m in M", where p(m) is the probability of m occurring, which may be 1/16 if all 16 mappings in M are equally likely to occur, but we place no such restriction on the distribution.

Next, let W = E[s*t|Q,R] + E[s*t|Q,R'] + E[s*t|Q',R] - E[s*t|Q',R'], which decomposes into "sum p(m)*(m(Q)*m(R) + m(Q)*m(R') + m(Q')*m(R) - m(Q')*m(R')) for all m in M". We know that any expression S of the form qr + qr' + q'r - q'r', where q, q', r, and r' are real numbers in the closed interval [-1, 1], must be a real number in the closed interval [-2,2] due to algebra. (Explanation: qr + qr' + q'r - q'r' is linear in all four variables, in other words, the partial derivative of S with respect to any single variable is an expression lacking that same variable; so, S must take on its maximum and minimum values at the corners of its domain. Thus, some integer inputs q, q', r and r' each in { -1 , +1 } must yield the expression's min/max bounds. Either employing further algebra or applying brute-force on all the 16 possible integer inputs, we find -2 and 2 are the bounds.) Therefore, W is bounded by "sum p(m)*Sm for all m in M" where each S is an unknown in the interval [-2,2]. Therefore, W is in the interval [-2,2]. (Explanation: p(m) are weights that sum to 1, and a weighted average of values in an interval must yield a result in the same interval.) This inequality, -2 ≤ E[s*t|Q,R] + E[s*t|Q,R'] + E[s*t|Q',R] - E[s*t|Q',R'] ≤ 2, is the BCHSH inequality, which gives us a formal mathematical constraint when believing spin is determined by local hidden variables and not a distant actor.

Bell's theorem concludes, this time finding a constraint for W using quantum theories. The wave function in quantum mechanics simplifies E[s*t|A,B] to cos(2B-2A) for any real angles A and B. By letting Q = 2pi/8, Q′ = 0, R = pi/8, and R′ = 3pi/8, then W simplifies to cos(-pi/4) + cos(pi/4) + cos(pi/4) - cos(3pi/4), which is approx. 0.707 + 0.707 + 0.707 - (-0.707) = 4 * 0.707 = 2.828, thereby violating BCHSH. This is perfect, it places contradictory mathematical constraints, between [-2,2] with local hidden variables and 2.828 with the wave function, on W.

Due to the law of large numbers, the running average, obtained by repeatedly rerunning experiments, rapidly converges to the expected value, meaning we can empirically measure E[s*t|A,B] for any angles A and B at arbitrarily high confidence levels. Current empirical evidence shows with high statistical confidence that the expected value is above 2.7, easily violating BCHSH and approaching 2.828 predicted by the wave function.



There is a caveat about Bell's theorem I had omitted from the above e-mail. A hidden variable could exist at the time of the universe's creation. Let's call this variable U, for universal simulator. If everything in the universe were proximate at this time of creation, then all particles may share this variable U, and any new particles may copy U whenever it comes across a U-carrying particle. Thus, all particles with high statistical confidence are U-carriers. Now, since U contains all information from the time of the universe's creation, then, if everything in our universe is pre-determined, we can know seemingly non-local information by merely querying the local hidden variable U, which, as a universal simulator, knows everything about our universe if events are pre-determined from initial conditions. Therefore, two entangled particles may "know" about each other from U and conspire to yield strange results.

Rather than a gaping loophole, this is a minor caveat because such a universe that conspires so maliciously to deceive us is akin to the ones imagined by philosophers of antiquity when they questioned whether the universe may have arisen only now, with every particle perfectly in place with the right momentum to trick us into believing there was any history at all when none existed. Were such celestial malice present, it would be impregnable to philosophy, science, and studies in general. Therefore, by reductio ad absurdum, if we wish to study, we must study the alternative, the scenario in which the universe is not so malicious.

Tuesday, April 22, 2008

Fun with Number Theory: Guessing Your Age

  1. Take the last digit of your age, multiply it by 2, then for every decade you've lived add 1 to it. Color this result Blue in your mind
  2. Think of your favorite digit, any digit at all, multiply it by 11.
  3. Add the last digit of your age to the previous result.
  4. For every decade you've lived subtract 1 from the previous result. This is your Red number.
Alright, now tell me your Blue and Red numbers and I will guess your age!
Blue Number:
Red Number:
age:

© 2008 by Thoreaulylazy. All Rights Reserved.


Spoilers:



Tuesday, July 04, 2006

Fee! Fie! Foe! Fum! I smell the blood of an English data miner

Oh, how I love engaging in quackery on the once noble field of statistics! I found in my gmail nest an email to myself, dated 18th of September, 2005, containing a dictionary analysis using the venerable turd-machine, Google. In the email were my dataset and its associated scatterplot piccies, as was a Google Ban™ screenshot which I shall forever keep with me as a data miner's evidentiary badge of honor. No one need fret, as a simple dhcp release and acquire is all that's required to set one's ipv4 address straight, and, in retrospect, to poison the address pool with a bum addy that cannot even 'oogle... Hmm...

Fig 1.  Google Ban™  screenshot through IE; 13,400 queries made in the span of approximately two hours

The datasets are partitioned by the host which I told Google to search within. I originally contrived an idea that I would somehow compare and contrast BBC, LiveJournal, Blogspot, and the general English corpus, but because of the nuissance posed by the wide fluctuation in the wordcount per webpage, I grew weary of trying to formulate insightful commentary that the data would support. The lazy bastard and quitter that I am, I shelved the idea of writing about the findings - or lack of findings - until this very epiphanal moment when writing about such a thing became the least boring in the long list of nauseatingly boring things I could be whiling away my time at.

Fig 2.  x-axis represents frequency rank, 1 being highest; y-axis represents frequency in units of webpages per billion in domain-specific corpora as measured by Google; dataset contains 3,500 randomly selected words from the 1913 Unabridged Merriam-Webster dictionary queried against each of 4 domains

The horizontal axis represents the ranking order of a word by descending frequency. That is, x=1 represents the most frequent word for each particular dataset, which may be "the" for one set, and may be "an" for another set; and x=383 represents the 383rd most frequent word for each particular dataset. The purpose of keeping things in descending order of frequency is to form a Zipf curve, which is visually smooth and tenders brownie points for allowing me to mention Zipf.

The vertical axis represents the search results returned by Google, multiplied by a coefficient that lets us pretend the most frequent word of any dataset has 1 billion search results. This multiplicative shifting was necessary to make the superimposition of datasets in a single plot less jarring. Unfortunately, such arithmetic jugglery creates jagged artefacts near the tail-end of the curves, as the words become less frequent. This unsightliness is the graphical pronouncement of integer search results multiplied by what was necessary to make the most frequent word show up as 1 billion. Under a log scale plot, there is little difference between y=1023 and y=1024, but quite a bit of difference between y=1 and y=2.

Fig 3.  unscaled version of Figure 2

A notable, consistently reappearing anomaly is a discontinuity in the curves. This discontinuity, what I would dub The Google Chasm were it not for my vanity insisting on it being called The Thoreaulylazy Plunge, shows a steep dropoff in search results once the search results reduce to a navigable amount, which is 1,000 if you ever care to try to navigate to further and further pages in the google resultset.


Fig 4.  asymptotic region requeried the following week to check for reproducibility; plausibly inflated results multiplicatively deflated through division by 10 (approximately)


As a logical being, I would exhaust all avenues of explanation before spouting off fervently in an accusing tone. That said, I am baselessly laying the blame squarely on a Google conspiracy to inflate search results by a factor of ten once the results are no longer navigable and hence no longer easily verifiable. At some point, I should access the Google APIs under the free academic license and find out once and for all what this crevasse is all about. A mosey down into yahoo-land and a repeat of this data mining escapade may also prove fruitful as ammunition in my Google conspiracy claim. Ah, who am I kidding, I'm not persistent enough to furnish any evidence.

highly frequent words in the english lexicon
universe bbc.co.uk blogspot.com livejournal.com
1 and and and not
2 home home that add
3 site help not help
4 information policy home and
5 that skip they site
6 help not no ask
7 policy that site information
8 not e don lost
9 e responsible see yahoo
10 see site them e
11 no see very press
12 program they today policy
13 press edition old legal
14 they related big that
15 related northern every every
16 science watch place kind
17 them pictures help no
18 today science found don
19 network no though anyone
20 skip them photo looking
-5 direfrench gristmill galopade meruit
-4 vorspielgerman congeries exaltee appurtenant
-3 grstorgegr enthymeme conferva inexpugnable
-2 slipstickcoll nosce pegomancy livraison
-1 metrongr federalize solecize chronogram

Monday, October 24, 2005

number theory: ah, those funny li'l modulos

There were only a handful of tricks with numbers that the typical elementary school student would be inculcated with during those years I enjoyed (well, attempted to enjoy) primary education. To cast the bleak bleaker, the divisibility test by 3 is often the only ubiquitously taught trick possessing some semblance of novelty. In case anyone wondered as a kid but failed to discover the reason, modulus arithmetic is the preferred foundation with which one derives efficient divisibility tests. For some reason still unbeknownst to me, I decided to raise this ol' topic and began concocting digit-wise divisibility tests for the primes 7 and 11. I urge readers to please have ready several pitchers of water, since what follows is fairly dry...

When X is a string (of some radix) of numerical value Σ(X[k]*radixk)∀k∈0..X.length-1⊆Z, then (int)X ≡ 0 mod Y iff: Σ(X[k] * cycle[k])∀k∈0..X.length⊆Z ≡ 0 mod Y, where cycle is the repeating series { radixk mod Y }.

Let's pick a radix we all know and love, base10:
Let Dn = { 100, 101, 102, ... }
Dn mod 3 = { 1, 1, 1, ... } = Cycle (1). Since cycle[k] = 1, we get the age-old mantra "X is divisible by 3 iff the sum of digits is divisible by 3".

But let's look beyond 3,
Dn mod 7 = cycle ( 1, 3, 3*3≡9≡2, 2*3≡6≡-1 ..) = cycle (1, 3, 2, -1, -3, -2)
Dn mod 11 = cycle (1, 10≡-1 ..) = cycle (1,-1)

Interestingly, but with less power than an iff relationship, since lcm(2,6) = 6 and both Dn mod 7 and Dn mod 11 have half-cycles (-1 as an element), then X ≡ 0 mod 77 if X in decimal form can be grouped into 3digit strings, where every other 3digit string is marked red and those in between are marked blue, and the { reds } minus { blues } = Null.

For exampe, let red = { 123, 444, 812, 912, 083, 948, 020, 436 }
Thus, blue = { 123, 444, 812, 912, 083, 948, 020, 436 }
Let shuffle(blue) = { 444, 948, 912, 436, 083, 812, 123, 020 }
Then interlace(red, shuffle(blue)) = 123444 444948 812912 912436 083083 948812 020123 436020. This now yields a base10 string which is divisible by 77 (and obviously also by 2, 5, 7, 11 and all the other factors of their multiple 770): 123444444948812912912436083083948812020123436020

Too big a number to verify with a pocket-calculator? Here's an easy multiple of 77 to generate: 001001 = 1,001. A pocket calculator can verify it is 13*77.

But then, what's the ratio of densities between these simple multiples and the full set of multiples? In the case of 77, cardinality of the full set of multiples for base10 string X is (1/77) * (10X.length). Cardinality for the parlor-trick partial set is:
Let redi = ∪(X[6*i+0..6*i+2])
Let bluei = ∪(X[6*i+3..6*i+5])
Let g = X.length/6
(Π(φ(∃unique j s.t. bluei = redj))∀i∈0..g-1⊆Z) * (10X.length)
= (Π(1 - ((103-1)/103)k)∀k∈1..g⊆Z) * (10X.length)
= (Π(1 - 0.999k)∀k∈1..g⊆Z) * (10X.length)
limX.length→∞( Π(1 - 0.999k)∀k∈1..g⊆Z ) ≈ 7.4210E-713

Thus, the ratio of the two is (1/77) : 7.4210E-713 ⇒ 1 : 5.7142E-711. In other words, the parlor trick while being a very good way to generate multiples is a rather improbable way to test for divisibility by 77. The only rigorous divisibility test is the one where iff is ensured instead of the comparatively impotent if.

Tuesday, September 13, 2005

neither 42 nor 47 are interesting

Infected with an unfortunate meme of geekdom, I've grown unnaturally sensitive toward hearing the numbers 42 and 47. These numbers are almost mythical in nature, 42 for possibly being "the answer to the ultimate question", and 47 for supposedly having an anomalously high frequency of use in daily affair.

In a bid to cure myself of susceptibility to mind control by those who'd reap the awesome power of these two numbers, and under the premise that few enough netizens are infected with these memes to constitute any significant alteration to these numbers' use — a premise I'm sure we can all agree upon given the growing ubiquity of non-geeks online — I thought I'd query Google, our very own 'Deep Thought', on what it thinks about single and double-digit numbers. Here are the results:
x-axis represents the number queried, y-axis represents matched documents in millions under Google's Sept 13, 2005 corpus

Thankfully, there existed no easily roused DoS-prevention logic and so my perl script was sportingly allowed to run unimpeded.

Observing the data, we see that lower magnitude numbers reign supreme as expected, followed by their products by 10 and to a lesser extent theirs by 5. Our odd decision to make seconds and minutes base-60 probably encouraged the high popularity of 30. Shopkeepers who price everything one or two cents shy of a whole buck are the likely reason for the spikes at 98 and 99. Curse their financial legerdemain! I snuck in 100 despite it being triple-digit to show my solidarity with our love of base-10. If you must know, I originally had 0 as well but it made the scatterplot rather messy with its frequency being between those of 12 and 13. Notably apparent from the chart, there is nothing strikingly special about 42 or 47.

So there you have it, confirmation that the '80s really were boring, and incontrovertible proof that neither 42 nor 47 are in any way more spectacular than other numbers in the hearts and minds of sane, normal people. The atypically higher frequencies of 44 and 64 do on the other hand raise new questions...

After all this faux statistical analysis of 'cult-figures', it is perhaps prudent to reflect on the sagacious words of Homer J. Simpson.
1F09, 1/6/94 Homer the Vigilante
Kent: Mr. Simpson, how do you respond to the charges that petty vandalism such as graffiti is down eighty percent, while heavy sack-beatings are up a shocking 900%?
Homer: Aw, people can come up with statistics to prove anything, Kent. Forty percent of all people know that.

His eminence's eloquent words denouncing statistically based reasoning. Amen, Mr. Simpson, amen.

Tuesday, August 09, 2005

Robert Blackwill, Dictionaries, and p[a-z]*pot(ate|ent)

Despite a gleaming mention in the subject, Robert Blackwill has very little to do with this particular journal entry. In fact, he is merely part of the ambient noise surrounding a slippery word that tantalizingly evaded me. It's quite odd how, when focusing so intently on recollecting one thing, loosely related items emerge into thought and, though unsolicited, they emerge with unbridled clarity and brilliance.

I began explaining to a friend how I hate synonyms and prefer words which carry a larger concept - a purposeful word useful in reducing the amount of time to convey an idea. At that moment, I caught the shadow of a word which would be of prime example; unfortunately, anything more than the shadow eluded me. Lilipote? No, no. What was it? I struggled to think. I knew its definition is the use of a euphemism or understatement for emphasis. I started rattling off to my friend whatever I could think to see if the word would come to me. If two people are berated by an extremely mean officer, one can later use hyperbole to tell the other "that's the worst officer ever in history!" One can also use sarcasm to tell the other "he sure is a nice guy." One can also use a .. er .. lilipote to tell the other "he's not an extremely nice guy." Tragically, I was in the car without internet access. I asked my friend over the phone to google "lilipote" and although I was expecting no results, it was disheartening to be confronted with that reality. It did make for a decent segue into Lilliputs and from there the conversion meandered to other things of interest, such as the movie about a 40yr old male virgin, and away from this mess of grappling with quarter-life senility.

To ensure no one's spirits have risen in thinking this blog will delve into the interesting parts of that conversation, I should make clear that interesting movie-related discussion will never be in any blog I write. Interesting stories can easily be told to people, real-time; it's the uninteresting things which require the occasional straggler who resorts to reading blogs, having tired of hours of solitaire.exe and an additional hour of repeatedly dragging rectangles on her desktop to see icons highlight. Right, now that we've squared that away.. Once I arrived home and the conversation ended, I managed to remember 'litotes' was the word I likely was looking for, and a quick validation on the web vindicated my belated answer. Naturally, just before feeling the burden lifted, I remembered once knowing a word used in a many years old NY Times passage about then ambassador Robert Blackwill being autocratic and sinking his staff into depression by incessantly denigrating them. I knew the word meant something along the lines of being given authority or power by appointment; other than that, I only knew it began with something like "pen" and ended with something like "potate" or "potent" which is sadly insufficient to query a web dictionary, or a printed one for that matter.

Cursed word-based queries, grep would solve this in seconds, I muttered. Wait, yes, grep will solve this in seconds! After some longer than anticipated tar.gz hunting, I seemed to find numerous providers of the 1913 Webster's unabridged English dictionary, which has fortunately been released into public domain. With some tr, sed, and sort -u, I had a nice greppable txt file revealing 'plenipotentiary', carrying as a noun the meaning "a diplomatic agent, such as an ambassador, fully authorized to represent his or her government" and as an adjective "invested with or conferring full powers." Aha, excellent. If only there had been a web service to do this generally! Are you listening, Google? I seriously need a mind-augmenting persistent cache to prevent memory loss. As an aside, as if this whole posting wasn't a giant one, a month and a half ago I managed to utter "Al Grove" in place of "Carl Rove", seamlessly morphing him with a certain former vice president and presidential candidate I'll leave unnamed. Needless to say, that gaffe instantly clinched an unofficial debate victory for my heckling opponent.