The best kittens, technology, and video games blog in the world.

Wednesday, June 09, 2010

Best sources of DHA omega-3 essential fatty acid

Axolotl by Ethan Hein from flickr (CC-NC-SA)

This is surprisingly difficult to find out, so I decided to share the results with everyone. But first, background:
  • Animals need omega-3 and omega-6 essential fatty acids
  • The same enzymes are used for omega-3 and omega-6 processing, so too much of one will interfere with the other. People used to have diets with about 1:1 omega-3:omega-6. Today ratio is more like 1:20, and that little omega-3 we eat is mostly ALA.
  • Three most important omega-3 are ALA (18 carbons), EPA (20 carbons), and DHA (22 carbons).
  • Brains are made largely out of DHA.
  • Land plants produce no EPA / DHA. None whatsoever. Zero.
  • Some land plants produce adequate amounts of ALA, but even this is uncommon.
  • Algae produce quite a lot of EPA / DHA.
  • Animals including humans can convert ALA to EPA, and then DHA, but this is a painfully slow and inefficient process; and over-saturation of omega-6 and many conditions interfere with even that much.
Based on this some people believe it would be wise to try to increase amount of omega-3 fatty acids in diets. Hard evidence is rather lacking, this is however to be expected as hard evidence of anything about diet is essentially nil. It's almost only short term studies of crappy proxies, and there are millions of reasons why this is just wrong. Anyway, concerning supplementation:
  • All mixed omega-3/omega-6/omega-9 supplements are waste of money - you're eating too much omega-6 already, and you can make as much omega-9 as you wish yourself.
  • You don't want generic "omega-3" supplements - most of these are ALA, which is of very limited use. You want DHA. At worst EPA. ALA is little more than filler, it's not bad for you but it's less relevant, and much easier to get via normal diet anyway.
  • If you're surprised why it's so hard to get DHA, this is possibly highly relevant.
Hopefully now you see why I'm measuring DHA, not anything else. And to keep science proper what I'm interested in is "% of calories coming from DHA", not "grams of DHA per portion" or anything like it. Portions are whatever manufacturer says they are, "per 100g" measures mostly tell you how much water foods have, and only "% of calories from" measures the right thing.

Escher Symmetry by Pieter Musterd from flickr (CC-NC-ND)

Best Sources of DHA


I digged through USDA National Nutrient Database, and this is what I found.
  • The data only contains "food" not supplement pills and such. These will of course contain highest concentrations. Rarely eaten foods like dolphin meat are not in the database, so I have no idea how nutritious dolphin sashimi would be.
  • The best source of DHA is unsurprisingly - fish oil. Salmon oil is 18.2% DHA and 34.2% omega-3 altogether; other fish oils are pretty good too, but not so much. Other oils like cod liver, sardine, and menhaden are 8.5%-10.9% DHA, 18.8%-26.6% total omega-3. Herring oil is less impressive 4.2%/11.1%. Fish oil also contains most mercury and other poisoning, so enjoy that. Once ultra-refined you won't need to worry about poisoning but it's more supplementation than food.
  • The second best source is caviar/roe, with 7.7%-13.6% DHA, and 13.7%-24.2% total omega-3. Might be expensive to turn it into a major part of your diet.
  • The third best source is seal oil with 6.5-12% DHA, 14.0-27.7% omega-3. Let's see if you can buy some legally outside Canada. So far we're totally out of luck.
  • Finally something more useful. The fourth best source is salmon. There's wide range of nutritious value from 2%/4% to 8.9%/13.7% - depending on where they're caught and what's their diet. Fortunately there doesn't seem to be a big difference between wild and farmed salmon, so either will work. Unfortunately the same mercury poisoning problem applies as to fish oil - poisonous substances are stored in oil so the oilier (and more useful for us) the fish the more toxic, and there is no way to escape that.
  • Fifth best source is mackerel. Like salmon, it can have as little as 1.5%/2.8% or as much as 8.7%/14.7%.
USDA says the next best source is dried parsley leaves at 7.3%/9.9% - which is most likely a massive measurement error, as there are no other plants anywhere, and it really makes little sense.

Old friend by JennyHuang from flickr (CC-BY)


Other than that, it's fish, fish, fish, mollusks, jellyfish, crustaceans, and more fish. My hopes for finding something that's not fish are getting slimmer and slimmer, so I'm just going to skip all of them now (you can probably see the pattern), and only focus on things which are not fish/seafood.
  • Brains. Beef/lamb/pork brains have 3%-5.4% DHA and 4.4%-7.7% total omega-3. Not surprisingly, as that's what animals primary need DHA for. And we simply throw away this most nutritious part.
  • Whale oil - 3.9%/8.3%. Whale meat on the other hand is pretty useless at 0.2%/0.5%. Not that you'll find much of either at the nearest supermarket. It's technically not a fish. Anyway, between brains and ocean creatures we pretty much ran out of good sources, the next source is:
  • Roasted squirrel - and that at mere 0.5%/0.6%
  • Chicken can be anywhere from 0.1%/0.2% to 0.5%/0.8%
  • Egg yolk - 0.3%-0.4% DHA, 0.4%-0.7% total.
  • Whole egg - 0.2%-0.3% DHA, 0.2%-0.8% total. Pretty much all of that in yolk.
  • We're long past useful concentrations anyway, so I won't be listing them. Next on the list are caribou, green turtles, turkeys, lamb kidneys, frog legs, guineahen, lamb hearts, squab/pigeon, pork livers, bear, raccoon, pork lung (and other beef/lamb/pork offal), and emu. Pork/beef/lamb meat doesn't register other than as rounding error.
So to summarize:
  • Get supplements;
  • Or eat a lot of fish and other seafood;
  • Or eat ridiculous amount of poultry, eggs, and game meat;
  • Or you're fucked.
There's no way to get enough DHA in anything resembling standard Western diet. Simply no way. Fruits and vegetables contain none, even "organic" ones. Actually fast food contains more as it often uses eggs and poultry as ingredients, but that's still not that useful.

One more long-term option would be to genetically engineer some common oily plants like soy or canola to produce some EPA/DHA - even if humans wouldn't eat them, if they're used as animal feed, we'd benefit indirectly quite a lot.

Actually someone already genetically engineered pigs to produce 4x omega-3 fatty acids, including 2x DHA, which sounds like the most urgent reminder that we need GMO now, and Luddites should not be in charge of policy.

EDIT: GM soybeans with more omega-3 (this most likely means ALA) and higher stability so they don't need partial hydrogenation has just been approved in US. If people moved from usual partially hydrogenated soybean oil to that it would be a massive health benefit. Of course our Euro-Luddites + CAP-paid farmer lobby coalition will probably ban it until long after "America vs Europe" picture get reversed. It's already much closer than you think.

Wednesday, June 02, 2010

What is Internet good for?

im in ur tube, blockin ur internets by the boy on the bike from flickr (CC-NC-SA)

Here a quick poll, pick the most fitting answer:
  • I spend too much time on the Internet, and I'm aware of this
  • I spend too much time on the Internet, I'm deluding myself about it
There's no need to bother with the third option, as people who don't use too much Internet are extremely unlikely to ever read this blog.

But in the spirit of "What have the Romans ever done for us?" - what good is Internet really for? Is it worth all the time we spend on it?

For the last two weeks in the spirit of Alicorn's luminosity - by the way definitely read that linked article, her first attempt was at that was a major tl;dr but these "seven shiny stories" are so short and insightful you're probably better off spending some time reading them than whatever else you're typically wasting your time on Internet, and it promoted her to my second favourite lesswrong writer after Eliezer.

So as I was saying before I interrupted myself, for the last two weeks I've been making a log of my daily activities and how much satisfaction I actually got from a given day.

Now many people who have learned economics 101 and are treating it too seriously think that whatever we're doing must by definition be the things we most enjoy, our claims to the contrary notwithstanding - so people who say the want to get thinner, but eat fuckloads of pizza actually prefer eating pizza and being fat to not eating pizza and being thin. This point of view is usually something worth considering - people usually say they want things they feel they're "supposed to want". On the other hand, the amount of evidence that we don't do what's best for us is ridiculously overwhelming.

A very very short list of such examples would include:
  • hyperbolic discounting - there's only one "mathematically consistent" way of treating values over time (exponential discounting), and the evidence is completely unambiguous that we're not doing so.
  • rodent experiments show also rather unambiguously that brains have separate systems for "wanting something" and "liking something". They are of course connected, so other things being equal if we like something more we will probably want it more - but this influence is far far less than total identity assumed by economics 101. Once you're aware that people might "want" things they don't really "like", and "not want" things they "like" (by the way - I'm using the words "want" and "like" rather vaguely - natural languages are spectacularly bad when analyzing humans - quite surprising actually as they seem to have been originally developed largely for social use) the entire utilitarian / consequentialist framework of analysis collapses.
  • Happiness research in spite of all its ambiguity at least shows that simple models of happiness are plain wrong. If you have time to waste on the Internet - TED is filled with good talks about happiness, not just the one linked.
  • Even disregarding these, you'd need to have ridiculously good information on yourself to decide what's the best thing to do. And it would be a massive understatement to say that we don't have it. And it's really really difficult to measure yourself. The idea behind "living luminously" is that increasing your self-awareness of your own mental state even somewhat might lead to highly positive results.
  • and many many more

Economics 101

By the way - and if you don't enjoy how I keep straying away from whatever is my main subject all the time, you probably shouldn't be reading this blog - I'm in no way disparaging "economics 101" thinking. I find it really sad that virtually every single person in the world pretty much belongs to these two classes:
  • People who don't understand Economics 101. They fail to get even such basic notions like comparative advantage or externalities.
  • People who get Economics 101 - and take it far too seriously. If you even briefly look at assumptions behind all its theorems, none of them is even approximately true in the real world. Very often you get lucky and this toolkit lets you predict things about the real world decently enough, but this is about as often not the case.

This is not to say Economics 101 should be thrown away. All models are wrong, some are useful. Or from a closely related perspective - all abstractions leak. If you don't perform sanity checks, and blindly trust everything the models tell you, they will lead you far astray (you could always argue that it's not models' fault, it's fault of the way you're applying them - but this is a purely theoretical distinction).

So for example in theory economics 101 models say that laissez-faire international trade policy should outperform any kind of intervention, but in practice countries which practice export subsidies by currency manipulation like the East Asia are better off than those with less laissez-faire trade policy like Europe, which in turn are better off than those that try to follow the route of import substitution via high tariffs like Latin America.

By the way if you find this curious, the standard answer to this puzzle is that benefits of economy of scale overwhelm loses due to comparative advantage - pumping subsidies into narrow range of related industries lets your country specialize in those, and import everything else - while limiting imports of wide range of products to protect diverse local industries means they will all be small and weak. As economics 101 completely ignores the dominant factor of scale advantages, focusing on an undoubtedly real but less important factor instead.

Similar misapplication of economics 101 says that minimum wage laws invariably increase unemployment. This was universally believed by nearly all economists a few decades ago according to some surveys cited by Wikipedia, and holding such belief became almost the canonical way of signaling that you're "economically savvy". And not surprisingly it turned out to be false, all research showing either no effect whatsoever, or effect that is really tiny compared to the huge increase in well-being of the working poor. The economists have finally figured that out, and they're more or less evenly divided on the question - even those standing against the minimum wage laws typically holding much more nuanced views - and yet some naive hard-liners still hold this as a measure of "economic enlightenment".
I'm in your PC, stealing your internets by the boy on the bike from flickr (CC-NC-SA)

My logs


Getting back to the subject, for each day in a bit over the last two weeks I recorded my activities, and some measures of satisfaction. The logs weren't terribly detailed, and "recording one's satisfaction" is exactly the kind of thing which would never get published in any reputable peer-reviewed journal. That's not to say that it's useless - "the official way of doing science" has led to many important discoveries, but it's for many questions it's been rather impotent, and it's important that people try different ways of finding things out too, if for nothing else then to fill in the blind spots of the mainstream science.

So my logs, of however dubious methodology they are, seem to point to the following correlations. I won't even bother pretending to have any "statistical significance" in all that of course, even forgetting about small sample of just two weeks measuring "statistical significance" necessarily assumes that samples are essentially independent, and they're nothing like that. The list is:
  • Physical exercise of all kinds - definitely positive - this isn't really surprising, as this is something that's highly enjoyable once started, but it takes effort to begin, so the hyperbolic discounting excuse applies
  • With video games it's mixed. First person shooter games like online Modern Warfare 2 have positive correlations, but Total War games negatively correlate with my end-of-the-day satisfaction, even though they're not really frustrating or anything most of the time.
  • Cleaning up my GTD system - definitely positive
  • Being productive at work, and in general getting done things I want to get done, especially the long postponed ones - definitely positive
  • Reduction in caffeine consumption - mildly negative, but that doesn't really imply anything about long term effects of different levels of caffeine, and it wasn't even as bad as I expected
  • Watching TV series, and reading books - mildly negative; this might be a false result, or the effect might be real but minor, in any case - there's no reason to do much more of those
  • And the largest and rather surprising correlation - spending more time online has a huge negative correlation with my satisfaction levels
Now there standard economics 101 disclaimer applies - these measures are necessarily marginal - so finding out that I'm better off exercising more and Internetting less implies only that I would be better off exercising a bit more than now, and Internetting a bit less than now, and it's more likely than at some point increasing amount of exercise and decreasing Internet use will make me worse off.

It's an interesting find that I seem to actually enjoy my work - I should probably put that one in my CV for future reference. And it's nice to get some insight on what kinds of recreational activities work better for me than others (due to small sample size less repeatable events like those involving interaction with other people not included). But the big find is that Internet is bad for me, and let's focus on that.

What is Internet good for?


Do you remember what life was before the Internet? It was horrible! We had to copy games from friends on stacks of floppies instead of just bittorrenting them! But really, what good is Internet for?
  • Email and IM are always far superior means of communication than phone calls (I really hate those, they should all die in fire); and are so much faster and easier than driving all the way to meet someone in person that they usually win, even if the throughput is somewhat less.
  • There's shopping - at which Internet really excels, most of the time. At least when you know exactly what you want, otherwise not so much.
  • There is information - but I'm far from happy about it. For some kinds of thing that you want to find, if they follow "keyword keyword of keyword" pattern, you can usually google or bing it out in seconds. Otherwise, all search engines become nearly useless, even if this information is somewhere. And it requires a lot of knowledge to turn a problem into unique keywords - very often you know little more than "X doesn't work", or "I'm not happy about Y", and search engines won't help you with those at all. This assumes information is even online in the first place - as very often it's not - it's really sad how nearly all research papers have been successfully pay-walled. It seems that "information wants to be free" only when the information in question is something on the top 100 bestsellers list of one kind or another, and not to the long tail which contains most of the real value.
  • There's Google Maps and similar sites, which are far superior to paper maps.
  • There are some funny things online - but they're swamped by such amounts of unfunny repetitive material that I really doubt Internet is even good for that. I dare you, go to let's say /r/funny on reddit - which is supposedly about the most recent funny stuff online (or at least reddittors seem to believe they find stuff first, and everyone else copies stuff from them) - and how many things you'll find there that will make you laugh, and are not nearly ancient? And it's the same on nearly every other place which is supposedly filled with funny stuff. People keep going there because occasionally something good turns out, but it's so rare it's probably not worth it.
  • There are news - and again the flood problem applies. I'm yet to find any RSS feed with only important news. Everyone seems to believe the right way to do news online is to just throw 10+ trivia items a day - Obama said something, one minor celebrity divorced another minor celebrity, stock prices decreased somewhere, IDF shot a few unarmed civilians somewhere else - as if knowing things like these made you better off in any way. The choice is to either get ridiculous amount of political trivia, or just ignore the news altogether.
  • There are all kinds of social networking sites - and I'm increasingly doubting their value. I have a Facebook account (and accounts on some other sites) and I might even use it occasionally, but I don't see that my life would be that much worse if Facebook and the rest didn't exist.
  • There are blogs, wikis, and similar places where you can contribute your knowledge, which can give tremendous amount of satisfaction to the contributor. If you add them all up, they provide a lot of value for readers as long as search engines manage to find a relevant blog post or wiki article, which is always in doubt. On the other hand, I have serious doubts about reading everything on a blog, or unfocused browsing on wikis. I haven't yet seen a blog which had consistently good posts - much less consistently good and relevant to my interests.
  • There are online video games, for some things playing with people is more fun than playing against computer.

These seem to cover the main points. And what's obvious is that the most valuable online activities - email, highly focused search, shopping, maps - take rather little time. On the other hand, the ones that take a lot of time - like all the reddits, forums, social sites, wikis, blogs, etc. - don't provide that terribly much value per time spent.


Unless you actually measure how you use internet, it's really easy to overestimate how much value spending time on it gives you - as you're far more likely to remember the high points which didn't really take long - as opposed to relatively pointless activities which took most of your online time. Human memory just works like that.


This distinction would be pointless if it was impossible to make a distinction between the two - if reduction in the bad kind of internet use required essentially proportional reduction in the good kind. Fortunately it seems to me that this is fairly straightforward - good things like email (this of course assuming you have a working spam filter / and all mailing lists etc. go somewhere else than your inbox), shopping, maps, directed search - have different entry points than less useful things like social sites, wikis, news, and funnies etc. Sometime you'll look for something specific and in the process accidentally fall into a wiki trap, but this shouldn't be too common with some self-awareness. Much more often you waste a ridiculous amount of time by wanting to "quickly check if there's anything good on X" and having hyperbolic discounting ("just one more link") turn that into a disaster.

It doesn't mean reducing your lolcat consumption to zero - only that you should force yourself to make an up-front decision "I'm start looking at funnies now, even though I know well enough it will probably take the next few hours" and having a realistic idea how good this time will be; instead of fooling yourself it will be quick and only filled with the good stuff, as seems typical now.

tl;dr - using internet only when you have clear goal, and not for vague "maybe there's something good" is good for you.

Saturday, May 22, 2010

Empire Total War mods - no walls, libertarians everywhere

The look of nobody home by Tjflex2 from flickr (CC-NC-ND)

Rome Total War was highly moddable. Medieval 2 Total War somewhat less as all data files were in obfuscated packages - but once you unpacked them it was as good as Rome. My M2TW mod is generated by a bunch of simple regexps, with which I can experiment as much as I like, turning features on and off and changing their magnitude in seconds.

I knew Empire Total War won't be as easy, but nothing prepared me for the pain I suffered. But I'll complain as I go.

Libertarians everywhere

First to see how modding works I wrote a very small mod - turning off taxes adds +10 happiness. Not quite literally, as it seems impossible to have anything triggered by zero taxes, so I gave all non-zero taxes extra -10 happiness, and every government type gets extra +10 happiness - net result being what I wanted, except it looks a bit silly in game.

Now why did I do so? The answer is my favourite "less micromanagement". I don't particularly having to babysit conquered provinces, chasing rebels around, and counting how many units I can move and how many I need to leave. This is simply not fun. So I decided that every province has a sizable population of armed libertarians who will gladly shoot every protester for me as long as I set taxes to zero. When taxes are not zero they blog climate change denial or something, I don't care.

This is all a stop-gap measure for the first few turns after conquest - if you ever want to get any taxes out of the province, as you usually do, you'll need to deal with taxpayers' happiness eventually. By the way zero taxes essentially means provinces loses money every turn, as it increases administration cost of every other province in your empire.

Happiness +10 is enough most of the time, but when you conquer someone's capital and have to face -30, on top of all industrialization, religion etc., you might need to deal with rebellions anyway. And as such regions usually bring a lot of money, you probably want to set higher taxes anyway.

If you want to tweak this bonus, use DBEditor for it.

No walls

And now the big mod - removing all walls. I'm not ideologically opposed to settlement fortifications - in fact my M2TW mod makes settlements more difficult to take. However:
  • ETW sieges are broken due to stupid AI and stupid pathfinding
  • Because they're broken, all ETW sieges follow just two boring scripts:

    • either: approach diagonally with infantry from 3 directions, bayonet charge everything;
    • or: approach diagonally with howitzers with carcass shot and a few line infantry units, once you start bombarding them AI will charge you one by one and you win without loses
    everything else fails as your units are too stupid to pathfind in more interesting strategies
  • After 1710 or so nearly every settlement outside Americas has city walls
  • And so after the first few years, 90% of battles is unbelievably boring
Field battles on the other hand are much more interesting. An unfortunate side effect of this mod is that while in vanilla town watch can use settlement fortifications to defend it from small enemy armies reasonably well, now it's completely useless. I can live with that.

Unfortunately, this has second order side effect of forcing you to be more aggressive. If before you'd be quite willing to have your armies further away from the front as town watch could serve as good enough first line of defense, now you need to have your armies closer to the borders, preferably on their side of it.

So how did I write this mod? It was truly painful. First, ETW has building features, and the fact that settlement fortifications building creates walls in the battle is supposedly encoded by such features. I ran into the first problem, as DBEditor is incapable of removing features - only adding or modifying them. To do such thing you need to go to PackFileManager, clone a table, make sure it's named exactly like the original one (so it will be overridden, not merged), and remove rows from that in DBEditor.

Except for some reason PackFileManager incorrectly mixes up slashes and backslashes, so I needed to fix the mod file from a hex editor.

Except it turned out in the end ETW ignores this feature, and simply hard-codes walls. So much for my effort.

It was plan B time. First I needed to remove all walls. And while db files are pain to edit it is nothing compared to the pain of editing startpos.esf which contains campaign information:
  • Start campaign, look at every single settlement on a map and write down which settlement has walls (as there's no search function in EsfEditor)
  • Manually find every such settlement in esf tree, and do all necessary changes (change True to False in one node, delete another node).
This is of course wholly incompatible with any other esf mod, like those which enable you to play minor factions.

After that I needed to make it impossible to build walls. I haven't figured out how to do that (I suspect if I remove walls completely the game will simply crash, and no other building in unbuildable) - so I made it require late technology of mass production, and take 999 turns to build. Sort of good enough.

And all this was possible only after a lot of effort of modding community members who created tools like DBEditor, EsfEditor, PackFileManager, documented what they could etc. This is borderline unmoddable, and I fully expect the next generation of Total War games after Napoleon (which is essentially extra scenario for ETW) to be completely impossible to mod. But hey, maybe they'll sell more DLC this way.

How to install

In case you want to install these mods, you need to:
  • Download starpos.esf and replace one in C:\Program Files (x86)\Steam\steamapps\common\empire total war\data\campaigns\main (or similar), first of course making a backup copy - this will remove starting walls. (small warning - it's a big file and the server I'm hosting it on is pretty slow; other files are tiny)
  • Download mod_no_walls.pack and put it in data directory together with all other packs - this will make walls unbuildable.
  • Optionally download mod_tax_break.pack and put it into data directory - this will give you +10 happiness bonus for no taxes.
  • Download ModManager and unpack it.
  • Start ModManager, select mods you want to use, and click Launch.

Greasemonkey and jQuery easier than ever

Douc Langur (Pygathrix nemaeus) by ucumari from flickr (CC-NC-ND)

Last year I wrote a short Greasemonkey tutorial, in which I explained how to use it with jQuery for some really simple scripts. Since then it became even easier, so here a few more useful scripts.

Since version 0.8 Greasemonkey has @require feature, in which your scripts can be made to depend on some external Javascript files - like jQuery. Unfortunately there are a few gotchas:
  • jQuery 1.4 doesn't work with Greasemonkey, you need to use jQuery 1.3 (or keep pestering them until it's fixed)
  • @require only works on installation time, you cannot use it with "New User Script..." feature, nor can you change @requires from the editor. If you simply use published scripts and don't write anything - you don't need to worry. If you write your own scripts you need to prepare something.user.js file somewhere, then open it from Firefox (copy&paste it's file path to Firefox URL bar, then click Install).

Show spoilers on tvtropes

Let's start with something really trivial. I don't care about spoilers. Maybe 1% of "spoilers" significantly diminish viewing pleasure, vast majority of them do not. I knew that Vader is Luke's father, that Snape killed Dumbledore, and that Titanic sunk before watching/reading, and it made it no less enjoyable.


On the other hand I'm quite annoyed that to read tvtropes I needed to constantly select "spoiler" text for everything to see it - but no more. This trivial script solves it entirely:


// ==UserScript==
// @name           Show all spoilers on tvtropes
// @namespace      http://t-a-w.blogspot.com/
// @include        http://tvtropes.org/*
// @require        http://ajax.googleapis.com/ajax/libs/jquery/1.3.2/jquery.min.js
// ==/UserScript==

$(".spoiler").removeClass('spoiler');

It's really simple - @name is unique name for the script, which should really be a very short description. We're not going to rely on @namespace at all, so put anything you feel like there.


@include is a pattern of URLs for which the script should be executed. @require is the jQuery library we're including. And thanks to jQuery, removing spoiler tags is really easy.

You can download this script here.


Always sort by seeders on The Pirate Bay

And extremely annoying thing about The Pirate Bay search engine is how it sorts results by its idea of "relevance" - usually giving you some dead ISO of three year old Ubuntu version when you obviously want the most recent one. And nearly always sorting by seeders is the right thing to do.

There is actually a GreaseMonkey script that claims to do exactly that, but it doesn't always work correctly, and the author decided to obfuscate the source for lulz. I'll have none of that, so I just wrote my own. Again, thanks to jQuery it's truly trivial.


// ==UserScript==
// @name           Always order by seed count
// @namespace      http://t-a-w.blogspot.com/
// @include        http://thepiratebay.org/*
// @require        http://ajax.googleapis.com/ajax/libs/jquery/1.3.2/jquery.min.js
// ==/UserScript==

$("input[name='orderby']").val(7);




You can download this script here.

Download Stewart and Colbert from bittorrent

Now a more complicated and perhaps more useful script. One of the recent developments on the Internet I hate the most are geographic restrictions - Americans can view whatever they like, for everyone else it's "Sorry, Videos are not currently available in your country" or other such nonsense.

Personally I don't care about licensing restrictions which led to this - I'm not going to accept development like that if I can help it. Fortunately in this case I can - all Stewart and Colbert is available on bittorrent.

The script does the following:
  • Find all dates on the page. They're not kind enough to use consistent class for those, I found at least 3, and perhaps I still missed a few.
  • Convert every date to from "Month DD, YYYY" to "YYYY MM DD" format, as this is what most torrents use. Half of the script is just doing that because Javascript doesn't have anything like Ruby's Date.parse.
  • Append link to bittorrent site next to the date. I'm using yourbittorrent here, as ThePirateBay is down due to overload far too often.
So every time someone on Reddit posts a "omg Colbert totally pwned Glenn Beck" link, I can actually see the pwnage now thanks to this script.


// ==UserScript==
// @name           View Stewart and Colbert on bittorrent
// @namespace      http://t-a-w.blogspot.com/
// @description    No more "Sorry, Videos are not currently available in your country"
// @include        http://www.thedailyshow.com/*
// @include        http://www.colbertnation.com/*
// @require        http://ajax.googleapis.com/ajax/libs/jquery/1.3.2/jquery.min.js
// ==/UserScript==

var month_names = {
  'January': '01',
  'February': '02',
  'March': '03',
  'April': '04',
  'May': '05',
  'June': '06',
  'July': '07',
  'August': '08',
  'September': '09',
  'October': '10',
  'November': '11',
  'December': '12'
};

$(".date, .airDate, .clipDate").each(function() {
  var txt = $(this).text();
  var m = txt.match(/(\S+)\s*(\d+),\s*(\d{4})/);
  var date = m[3] + "+" + month_names[m[1]] + "+" + m[2];
  var url;
  if(window.location.hostname == "www.thedailyshow.com") {
    url = 'http://www.yourbittorrent.com/?q=Daily+Show+' + date;
  } else {
    url = 'http://www.yourbittorrent.com/?q=Colbert+Report+' + date;
  }
  $(this).append("<div><a href='"+ url + "'>Download on bittorrent</a></div>");
});

You can download the script here.

See how easy it all was? Now start writing your own scripts and share them.

Thursday, May 20, 2010

Everybody Draws Muhammad Day and lessons in cultural relativism

May 20th is Everybody Draw Mohammed Day - a day when you too can join the defense of Free Speech by drawing a cartoon. Or if you really suck at drawing you can make a shop of Muhammad like me.



Prophet Muhammad, artist's conception
If you think I'm wrong, draw a better one

Lesson #1 - Principle of dissimilarity


Recently I have been convinced by some TTC audiobooks that Jesus might have probably existed as a historical person. You know what's the best argument for it? The criterion of dissimilarity.

It's not any pomo deconstructionism - just plain old historical analysis. The idea is that authors have an agenda, but life doesn't always agree with it. So every time the text admits to something that is clearly against the authors' agenda, it suggests it probably had some basis in reality - because they wouldn't make this up on purpose. A few examples from Jesus' life:
  • Crucifixion was embarrassing kind of death reserved for slaves, rebels, and other lowest status scum. If someone was making up the story they would have Jesus die in battle, or die some other "respectable kind of death". So actual Jesus was most likely actually crucified.
  • All the messy explanations how "everybody thought Jesus was born in middle-of-nowhere Nazareth but actually he was born in Bethlehem like King David" suggest that Jesus was probably born in Nazareth. If they were making this up, they'd skip the Nazareth story altogether.
  • Jesus was baptized by John the Baptist - being baptized by someone was signaling submission and inferiority to that person. This story is clearly embarrassing to Bible writers, they put words like "I should be baptized by you" in John the Baptist's mouth, and the last Gospel omits it altogether.
The sheer volume of such cases suggests that Bible is an "enhanced" story of some actual life, as opposed to being completely made up. There's no question that all the supernatural bits have been inserted, and it's widely believed that the real life stories were quite significantly massaged - but it does such a bad job at covering a lot of embarrassing parts like these that the hypothesis "Bible is vaguely based on a life of a real person, and these embarrassing stories were widely known so couldn't be ignored" is much more compatible with the evidence than "Bible is completely made up" - and as their Bayesian priors are not terribly different in the first place, historicity of Jesus can be reasonably believed in. The "Bible is real" story on the other hand is both against the priors and against the evidence, so let's not even go there.

Rambo Cat by Gerard Girbes from flickr (CC-NC-ND)

Lesson #2 - Why is Muhammad widely considered a pedophile?


The criterion of dissimilarity must be truly infuriating to the true believers - it essentially says that every time they like some part it's probably false, and every time they dislike some part it's probably true. By the way this applies to all historical texts, not just religious ones - parts of Commentarii de Bello Gallico that are overly sympathetic to Julius Caesar should be looked at with more suspicion than parts which talk about his failures, and so on.

So, what does it all have to do with Muhammad being a pedophile? Muslim texts clearly say that he was married to Aisha when she was 6, and fucked her when she was 9 year old.

Now how likely is it that this particular bit was made up? If you had a prophet who didn't fuck children, would you casually make up a story that he actually did? That'd go against all rules of proper writing. And it's not one isolated verse somewhere - as Wikipedia says "references to Aisha's age by early historians are frequent", and nobody questioned that back then.

So regardless of our believes, we can be fairly certain that
Muhammed fucked a 9 year old girl

And now some time for cultural relativism - this was considered fairly unremarkable back then. Our modern culture is obsessed about young people's sexuality and as a civilization we essentially lost the ability to propagate the species, but historically it was entirely normal for much younger people to marry, fuck, and have babies - biologically late teens are the optimal time to have children. In completely typical Ancient Rome - girls typically married in their mid teens, very soon after puberty. The idea was simple:
  • Regardless of societal constraints, many people will get sexually active once they hit puberty
  • Of people who get sexually active, many will get pregnant
  • Nearly everywhere except for modern times, it's really difficult to either provide for your kids without a husband, or get a husband if you already have kids
  • To avoid risking that, it's better to marry your daughters sooner

This was the baseline normal case for girls.

Western pedophile obsession (where all sexuality of people under 20 or so is a big taboo - this has little to do with what psychology calls "pedophilia") is a highly atypical case. 7th century Arabic children marriage, and fucking of prepubescent 9 year old girls was also rather atypical.

Not that it matters what's typical and what's not - abject poverty, illiteracy, and lack of broadband Internet are historical norm, and yet I much rather prefer atypical modern situation.

Anyway, what we see here is a conflict of values - Muhammed fucking a 9 year old was acceptable within his culture, and is not considered acceptable within ours - not even for most modern Muslim countries. Aisha most likely didn't mind getting fucked, and nobody was weirded out by this, in spite of modern fiction of "age of consent". She was most likely not scarred for life or anything like that - these laws exist primarily to make adult bigots feel good, not to "protect" children, and evidence is clear that plenty of teenagers (and an occasional pre-teen) have sex and enjoy it.



Lesson #3 - Freedom of speech


But that's not all. On top of one value dissonance about sex with pre-teens we have a second value dissonance about free speech and religious respect.

In Western culture since the Enlightenment, we have been very strongly attached to the idea that there is nothing that cannot be criticized. Yes, laws of different countries prohibit different kinds of speech (including United States, Supreme Court's favourite activity is making exceptions to the First Amendment and Elena Kagan doesn't seem any different here) - but every time such law is applied the media get freaked out and people feel highly uneasy about it. Even people who want to limit freedom of speech considerably universally consider it to be the default case - exceptions to be few and made only when "necessary".

These believes are not shared by many, apparently including modern Muslim societies. It seems that they approach Muhammed cartoons from another direction - that people's religious believes should be respected, and their holy figures shouldn't be mocked - freedom of speech being far lower in their order of priorities than that.

They don't even see the cartoons as a freedom of speech issue, just like 7th century Arabs didn't see fucking a 9 year old girl as a child abuse issue, and we don't see the cartoons as a blasphemy issue. Different cultures have different perspectives.

This basic cultural relativism does not mean that all cultures are equally wrong, or equally right. It would be intellectually dishonest to think that your culture is uniquely correct about things, and its beliefs are some sort of human universals - history shows that there are hardly any true human universals, and we might even get rid of death and taxes one day. But you are still free to follow your own culture's value system.

If you think that freedom of speech is valuable, and fucking prepubescent girls is creepy, mocking that and not giving a shit about angry Pakistanis is all fine.

In other words - enjoy the Everybody Draws Muhammad Day.

Monday, May 17, 2010

Very simple parallelization with Ruby

kitten geometry by tizzie from flickr (CC-NC-SA)


I hate waiting. And all too often when I tell computer to do stuff, the only reply I get is a progress bar with ETA far in the future.


Fortunately plenty of such things are highly parallelizable. For this post let's focus on a simple task of downloading every single Questionable Content strip ever released because QS has really simple URLs. It's real simple to do it serially but waiting time of about 23 minutes is killing me:

(1..1665).each{|i|
  system "wget", "-q", "http://www.questionablecontent.net/comics/#{i}.png"
}

Notice how I didn't even bother with shell, and went straight for a real programming language. Anyway, Ruby has very easy threading system - just use Thread.new{ ... } to spawn, Thread#join to join, avoid global mutable state, and you're done.

We can get it running in no time:

module Enumerable
  def in_parallel
    map{|x| Thread.new{ yield(x) } }.each{|t| t.join}
  end
end
 
(1..10).in_parallel{|i|
  system "wget", "-q", "http://www.questionablecontent.net/comics/#{i}.png"
}

The only thing that changed was replacement a one-liner middle-ware method, and replacement of Enumerable#each with Enumerable#in_parallel - so far really good.

On the other hand, we don't want to spawn 1665 threads all at once - neither our network connection nor QS's servers would appreciate that much. Let's write some code which would create 10 threads and keep them all busy instead.

Exception handling interlude

A very common problem with Ruby multi-threaded programming is that we don't want the program (or even thread) to die just because of some error - usually we want to capture, print, and otherwise ignore exceptions. Here Ruby fails hard - Exception has no method to generate user-friendly message of the kind it prints when exceptions fall out of the main loop. There are some gems that handle that, but here I'm rather unconcerned with this - and instead I'll just use a really simple catch-print-ignore method - which does not even bother printing backtrace.

def Exception.ignoring_exceptions
  begin
    yield
  rescue Exception => e
    STDERR.puts e.message
  end
end

Black Beauty by jakeprzespo from flickr (CC-BY)


Execute N tasks in parallel


How do we communicate with N threads? A simple and wrong way would be to create N queues, and feed them in a round-robin manner. Unfortunately unless tasks always take nearly the same amount of time (this is almost never even close to true) - this results in a lot of threads waiting aimlessly.

We need exactly one todo queue. Now it would be better if this queue had a size limit - but SizedQueue was giving me troubles, so I did it with the plainest possible Queue.

The main thread is really simple:

require 'thread'

module Enumerable
  def in_parallel_n(n)
    todo = Queue.new # Create queue
    ts = (1..n).map{ # Start threads
      Thread.new{
        # Do stuff in threads
      }
    }
    each{|x| todo << x} # Push things into queue 
    ts.each{|t| t.join} # Wait until threads finish
  end
end

Now the main question is how do we tell threads there are more tasks. Unfortunately queues have only two states:
  • Queue not empty - dequeue one element, do work, go back for more
  • Queue empty - sleep and wait for the queue
While we need three:
  • Queue not empty - dequeue one element, do work, go back for more
  • Queue empty, producer not finished - sleep and wait for the queue
  • Queue empty, producer finished - finish thread
Now I could go for a more complicated signaling solution - but that would either introduce global mutable state (hell no!), or require too much complexity - so I simply went for the simplest possible "flood queue with nils once you're done" way.

And because I don't want the program to break when object we're iterating over contains legitimate nils I'm passing whatever we get in a one-element array (Ruby probably has equivalent of Ocaml's 'a option / Haskells's Maybe monad etc. - but for this the simplest solution will work just fine).


require 'thread'

module Enumerable
  def in_parallel_n(n)
    todo = Queue.new
    ts = (1..n).map{
      Thread.new{
        while x = todo.deq
          Exception.ignoring_exceptions{ yield(x[0]) } 
        end
      }
    }
    each{|x| todo << [x]}
    n.times{ todo << nil }
    ts.each{|t| t.join}
  end
end


And now we can happily get the comics. It took 4 minutes 11 seconds instead of 23 minutes - 8x speedup!

(1..1665).in_parallel_n(10){|i|
  system "wget", "-q", "http://www.questionablecontent.net/comics/#{i}.png"
}

What about other comics?

Now unfortunately far too many comics don't have pretty URLs like that. You can always (well, nearly always) resolve to wget --mirror or find a greasemonkey script which makes the comics' website bearable. Fortunately quite often we can handle even somewhat more complicated URL schemes.

First, for comicses like sinfest, which use dates in URLs, it's only marginally more difficult:

require 'date'
(Date.parse('2000-01-17')..Date.parse('2010-05-17')).in_parallel_n(10){|d|
  system "wget", "-q", d.strftime("http://www.sinfest.net/comikaze/comics/%Y-%m-%d.gif")
}


3536 files, 7 minutes 52 second.

Because Ruby dates are first class objects and Range class works very sanely with them - it took nearly no effort to code that, other than trying to remember wtf were all those strftime %-codes (or you could use Date#year etc. and plain string interpolation - it would probably be better for your sanity actually).

And wget will quietly ignore proper 404s, so you don't need to wonder if they're updating on weekends too, or only on weekdays, or what.


For other comicses like xkcd you might need some hpricoting (and put alt text in metadata, then teach image viewer to read them? Or use imagemagick to put them into the image? choices, oh all the choices...). Or you could use This awesome Webcomics Reader Greasemonkey script. Unfortunately it only knows about a handful of the most popular ones.

Now this technique is applicable for a lot of uses other than downloading funnies - like handling S3 backups, encoding large number of MP3s and so on. Everything where it makes sense to run multiple threads in parallel, but not too insanely many will be really easy to do now thanks to Enumerable#in_parallel_n or similar code.

It works with both 1.8, and 1.9 equally well, in spite of their completely different threading systems.

Enjoy.

Tuesday, May 11, 2010

Steam is ripping you off

Online distribution of content should be cheap and easy. Compared to the complexities of making and moving around DVDs and such, moving bits is pretty much free. So it's a no-brained that games on Steam should be cheaper than the same games on DVDs, right? Of course that's not how it works.

These are top 10 Steam's best-selling games - this list is rather generous for Steam, as people would prefer to buy games on Steam which were better deal on Steam than on DVD, and would prefer to buy games on DVD which were better deal on DVD - so it's reasonable to expect comparing other games Steam would come out even worse.

As for DVD alternative, I wasn't looking too hard - just Amazon. I'm sure there are ways to get the games even more cheaply, but let's stick with the default online shop. Here are the prices, all Amazon's are with free shipping:

  • Call of Duty: Modern Warfare 2 - Amazon £29.65; Steam £39.99 - Steam costs 35% more
  • Battlefield: Bad Company 2 - Amazon £24.99; Steam £29.99 - Steam costs 20% more
  • Left 4 Dead 2 - Amazon £19.98; Steam £19.99 - essentially equal
  • Counter-Strike: Cource - Amazon £14.99; Steam £13.99 - Steam costs 7% less
  • Grand Theft Auto: Episodes from Liberty City - Amazon £14.99; Steam £19.99 - Steam costs 33% more
  • Civilization V (pre-order) - Amazon £30.99; Steam £29.99 - Steam costs 3% less
  • Mount & Blade: Warband - Amazon £16.97; Steam £24.99 - Steam costs 47% more
  • Borderlands - Amazon £11.78; Steam £19.99 - Steam costs 70% more
  • Just Cause 2 - Amazon £23.92; Steam £29.99 - Steam costs 25% more
In other words - median cost on Steam is about 25% higher than on Amazon. That's even though physical DVD costs quite some money to make, packaging, ship etc., while for Steam it's just pure profit.

And please don't even mention "Recommended Retail Price" - if you fall for this kind of blatant anchoring effect, there's no hope for you.

So to summarize: Steam games cost a lot more, and to make matters worse have no resell value, they're simply ripping you off.

PS. Amazon prices are for NEW games. These are not used titles. Anyone claiming otherwise has no idea what they're talking about.

PS2. Prices I gave are for UK (notice the £s). acct_rdt made a similar list for US prices, which shows exactly the same thing - US Amazon is cheaper than US Steam too, even though the difference is not as drastic (median 7%)
  • Call of Duty: Modern Warfare 2 - $37.80 new on Amazon, $59.99 on Steam
  • Bad Company 2 - $43.83 Amazon, $49.99 Steam
  • Left 4 Dead 2 - $29.99 (both)
  • Counter-Strike: Source - $19.99 (both)
  • Episodes from Liberty City - $27.99 Amazon, $29.99 Steam
  • Civilization V - $49.99 (both)
  • Mount & Blade: Warband - $24.53 Amazon, $29.99 Steam
  • Borderlands - $27.99 Amazon, $29.99 Steam
  • Just Cause 2 - $39.99 Amazon, $49.99 Steam

Greeks can pay, Greeks will pay

Roborovski Hamster Dwarf by cdrussorusso from flickr (CC-BY)

Most normal blogs' favourite countries to rant about is US, or Israel, or China, or something like that. Not on this blog - my number one target, at least for the last month - is Greece. And I'm really shocked that some of the otherwise sane people I know buy into the whole "can't pay, won't pay" nonsense, so I wrote this short post explaining how it all really works.

Basics of money lending


Let's get back to basics. Why would anybody lend anybody else their money? The most common reason is they want to get even more money back. Now this is in no way the only reason - people routinely lend money to their family members and friends, IMF lends money to countries facing collapse, USA and Soviet Union lent money to support whichever dictator seemed more aligned with their interest, Europe lends money to shit-poor countries out of pity, AIG lent itself taxpayers' money by bribes and threats, and CIA lent money to Osama bin Laden because they're idiots - but while these are all very important cases, most money in the world is lent for profit motive.

This profit-motivated lending is not only largest in volume - it's also most reliable. Political will can evaporate overnight - ask South Vietnam or Cuba if you don't believe me - but speculators can be reasonably counted on to keep lending as long as it makes business sense. Even in the middle of the latest recession and the Great Depression solvent borrowers had no trouble getting loans - the problem was sudden decrease in solvency, not disappearance of the profit motive.


Political and charitable lending is highly complicated and varied form case to case, but for-profit lending is quite straightforward from game theoretic point of view:
  • Lender decides to give borrower money or not
  • Then borrower decides to repay or not
  • Because borrower would much rather get money, and not repay it - no lender would be willing to lend anything without some pretty hard assurance of seeing their money back
  • On the other hand, as repayment might be impossible due to objective circumstances, borrowers would be unwilling to provide too much assurance
  • Level of assurance is based on balance of these two factors.

And what forms of assurance are available? For private individuals there's debtors' prison - a barbaric institution which is still used in child support and tax cases - not coincidentally both being the kind of "debt" which is incurred unwillingly, where the "borrower" has no leverage whatsoever.

Another assurance is forceful confiscation of property, and various kinds of physical abuse - these were routinely used as international debt collection in 19th century, see Egypt, and Haiti for examples - but these rarely happen these days.

Yes, lenders could try to sue Greece or another unwilling country in foreign or domestic courts, but it would be mostly a PR stunt, and they wouldn't even get enough to cover their legal costs. If a country doesn't want to pay, lenders cannot do shit.
Russian Dwarf Hamster Winter White by cdrussorusso from flickr (CC-BY)

Hamster showing what debtors' prison looks like

The final assurance


So what keeps borrowers repaying and lenders borrowing? It's lenders not being stupid. They know very well that if someone didn't honour their debts in the past, they're unlikely to do so in the future - so if you tell the banks to go fuck themselves, you saved yourself a huge pile of money, but you have no chance of getting any money from them in the future. 

When you think about it - it's a pretty weak assurance - it only works when the country is moderately-screwed-up:
  • If country's economic situation is truly screwed up - they won't be able to repay - so banks lose
  • If country's economic situation is bad but not horrible (like most countries now) - they cannot afford losing their credit lines - so they keep repaying - and banks win
  • If country's economic situation is some awesome they don't care about future loans - they can tell banks to go fuck themselves - and banks lose too!

Fortunately for banks, the last case is very difficult to achieve - not only you need zero deficit (before interest) now - something already very hard - you need to be certain of being able to keep this zero deficit essentially indefinitely. Recessions, commodity price fluctuations, wars, terrorist attacks, aging population, floods, tsunamis, volcanoes, epidemics, electing Conservatives, and plenty of other natural disasters can screw your budget essentially overnight - all sending you crawling back to the banks begging for money. And bankers will remember what happened to their old loans.

I cannot think of any country in the world which is financially healthy enough to pull that off. What makes it even less likely is that countries with most balanced budgets tend to have lowest debt levels and borrowing costs as a rule (debts are results of past performance, and past performance is the best predictor of future performance there is) - so for them repaying debts is nearly painless, and provides good insurance in case something screws their economies in the future.

Essentially the only case in which a country can tell the bankers to fuck off, and has debts high enough to make it worthwhile - is one which used to be horribly mismanaged for many decades, and then turned into one of the healthiest economies in the world essentially overnight.

Does it sound like Greece? If you believe so, I have some sovereign debt credit default swaps that you might be interested in.

Greeks simply cannot afford not paying. They can negotiate better terms of repayment, and get some help from EU and IMF, but essentially they will have to get their country in order - cut their bloated civil service and bloated military, cut - and repay their debts, because they simply cannot afford losing their credit lines. And so they will pay.

Friday, May 07, 2010

Audible review

dia mundial do rock by deadoll from flickr (CC-NC-SA)


Before the posting frequency either reaches singularity or crashes like Nick Clegg's hopes, here's a brief review of Audible.

I have listened to ridiculous amounts of audiobooks (a term I'm going to use for essentially any spoken word - whether an original material or a spoken version of a paper book) for many years now - a very long time before they became anything even vaguely approaching mainstream. And obviously I got most of them from Linux ISO download sites, not counting a very large number which were officially free as podcasts and such.

Then, my cat threw my iPod into a bathtub full of water, I bought a new and much better MP3 player, and with it came free subscription to Audible, which I actually extended for a few more paid months more out of curiosity how this world of paid audiobooks operates than anything else.

DRM


First - DRM is as bad as you'd expect. Devices cannot be activate on anything except Windows (the stories about iTunes being able to activate MP3 players are a lie). Then after I formatted the MP3 player, it lost activation and it was impossible to get it back. There was no error message, nothing. Activated successfully, still doesn't work. I had to email Audible for them to fix it... it was really one big pile of ridiculousness, all horror stories about DRM basically came true.

Does it at least DRM the books? Not at all. First, they left CD burning as an option, so anyone who is desperate enough can get virtual CDs and then rip them. And you can download Audible-ripping software which will happily hack into Audible drivers on any activated PC, and convert all your DRMed files into MP3s, FLACs, or whatever. Essentially they fuck over paid customers, and the pirates will get their warez anyway.

This is all good enough reason for not using Audible even if they didn't suck in other ways, but suck they do...

Speech synthesis


But before I get to the main dish of Audible suckiness, a brief interlude for state of speech synthesis. 60 years ago, when computing started, people believed that in no time computers will be as smart as humans. You know - 10 years, 20 tops. Which as it turned out in computing world means essentially "never".

I'm just amazed at how filled with wrong are predictions of people who expect rapid arrival of AI - like most users of the ironically named LessWrong community blog. How the fuck are we going to see AI go "foom" if during the last 60 years there has been essentially zero progress at:
  • Speech recognition
  • Handwriting (and most importantly hand-written math/diagrams/mindmaps) recognition
  • Speech synthesis
Yes, they expect a big foom - but better be worried in case this foomy future Kindle 2020 or whatnot will turn out to be unfriendly - and let's say will read the latest Twilight sequel in a very ironic voice...

The truth is - automated speech synthesis of anywhere near human quality is not coming anytime soon. It just isn't. Amazon knows it as well, while building gimmicky crappy speech synthesis into Kindle, they also purchased Audible for $300M whose business model totally relies on speech synthesis being crappy for decades to come! How is this $300M for a prediction market? Are you shorting Amazon stocks already? I didn't think so.

And while I'm at it, remember how one of my favourite crackpots - Raymond Kurzweil boldly predicted that by 2009 "[people] interact with their computers primarily by voice and by pointing with a device that looks like a pencil. Keyboards still exist but most textual language is created by speaking". Obviously this crackpottery never came true, I'll keep my Microsoft Natural Ergonomic Keyboard 4000 thank you very much. I doubt even Apple would be stupid enough to go into the whole voice thing - but then with iPad they sort of have proven that they can sell absolutely anything, and hipsters will buy it...

iCat by Sontra from flickr (CC-BY)


Why Audible sucks


And now the main part - why Audible sucks. To explain this think for a moment what an Internet shop is. Yes, it needs to have some stuff to sell, and some way to deliver it, and some way to accept payment, but primarily:
An Internet shop is primarily a search engine

That's right! Everything else is just an insignificant attachment to the main functionality of letting people find the stuff they want to buy. Go to Amazon - the archetypal Internet shop. What's there - a huge search engine. With user reviews, recommendations based on your previous purchases, and purchases of other users, and a huge wealth of information which helps you find things you might want to buy. Why nobody has copied Amazon yet? Well, there are some economies of scales involved, but dis-economies of scale exist as well. This is not the answer.

The primary reason is that nobody else has this massive database of data which helps people find what they want. Even if you had all the products Amazon has, at identical prices and terms of delivery, without user reviews and contextual recommendations your shop would be hopeless. And without even as much as Amazon search system? Forget about it.


By the way, I'm probably biased as most of my commercial career was in one way or another related to building custom search engines, just saying for the sake of full disclosure...

Anyway, Audible fails so hard at search it's not even funny. You go to category - and it's sorted by release date. How is this sane? You can sort by customer review - but most of their books have at most one review and there's no Bayesian filtering applied, so everything on the top will be audiobooks with a single 5-star review. Yes, it's all great that ONE person liked it, but it has extremely low predictive power on how good it actually is. About the only useful way to sort anything is by top selling.

So what do we see - title, author, price, star rating, a picture (wtf? these are audiobooks, these pictures are essentially just some random stock photos of no relationship with the product) - and that's it. And max displayed at once is 30 - but don't worry, it will change back to the default useless value of 10 results as soon as you change anything.

This is not Google! You don't want to go from search term out of their site in less than a second - you want to figure out what's actually any good! The logic of having smaller number of results so they're faster doesn't apply at all!

So how do you figure out if the book is any good at all? Well, there are very rarely any customer reviews (just a few star ratings is typical) and the ones there are are typically as useless as "Took me a little while to get into this book, but when I did I absolutely loved it! Have just downloaded the second book and can't wait to begin!". How does this provide any information whatsoever?

There is usually a short publisher-provided blurb telling what the book is about - and you can probably guess what I think about such publishers' honesty.

The only thing that's of any use at all - and without which the entire Audible website would be utterly useless - is the audiobook sample. Unfortunately it doesn't work terribly well. One thing you can figure out is that quality of recording tends to be top notch - unlike some other sources which compress audiobooks as if people were still downloading them on dial-up modems, audible handles quality properly. And you can also find out if you like the narrator or not. Which might be very highly relevant for their "erotic" section at least... (seriously, go there for some major lulz and/or facepalming..., but then maybe I'm not the target audience...).

What you'll rarely find out is if the book is actually any good, which is what we want to establish. Random pages from the middle of the book rarely tell you that. Even worse, I've seen quite a few books for which the random sample was actually legal blurb + acknowledgments from the first few pages... that definitely is some fascinating book, right !?


Unfortunately that's all your search options. Audible won't search inside books to figure out what they're about (speech recognition obviously being crap, but on the other hand - they have access to written original most of the time, so that's a weak excuse), and there's no other information on their website.

If you want to find out a good audiobook, you need to check every audiobook's written equivalent's page on Amazon - for reviews there. Somebody should write a Greasemonkey script for that... (this problem also affects UK Amazon to some extent - often US version has 10x as many reviews, but there's no link and the only way to get to those reviews is Googling... it gets me raging every time...)

How difficult would it be to link it up with the rest of Amazon? Seriously now.


There's also another issue - the choice on Audible is simply less than on "Linux ISOs download sites" - so when you want an audiobook it's very likely you won't find it there, even if it exists. The opposite - Audible having something which pirates don't also happens, but usually for more obscure titles.

This is actually an opportunity for them - sure, many people will torrent Dan Brown, but there's long tail to explore - except it is exactly this long tail which is it direst need of good search functionality. And so Audible fails.

So maybe it's a good idea to short them after all...

Thursday, May 06, 2010

Progress bar for Unix pipes

Don't Spill! by Ack Ook from flickr (CC-SA)


This is getting quite ridiculous, as it's the third post I'm writing in just one day - that's more than I typically wrote monthly over the last two years. This is what you get for finally getting organized - a big pile of 90%-finished stuff which while useless on their own can be very quickly be turned into something of high utility. OmniFocus and GTD are truly awesome. (unfortunately this means there are now two apps for Mac I care about - TextMate and OmniFocus - so switching away from Mac will be even hander)

The script I want to show you today solves one of the most severe problems with Unix pipes - lack of progress indicators. Normally you'd start a pipe, and until it actually finished what it was doing, you'd have no way of finding out if it's 1% done or 99% done, and how fast it is progressing.

There are some nasty hacks. Many times I used strace or looked inside /proc to figure out how much progress has been done - but these are painful waste of effort for something that should be builtin. Feedback is the basic principle of good UI design.

Anyway, here's the final solution to progress bar question:

#!/usr/bin/env ruby

STDERR.sync = true

$bytes = true
$max = nil
$count = 0

until ARGV.empty?
  case (arg = ARGV.shift)
  when '-l'
    $bytes = false
  when '-b'
    $bytes = true
  when /\A(\d+)([kmg]?)\Z/
    units = {'k'=>2**10, 'm'=>2**20, 'g'=>2**10, ''=>1}
    $max = $1.to_i * units[$2]
  else
    raise "Unrecognized argument: `#{arg}'"
  end
end

$max = STDIN.stat.size if $bytes and STDIN.stat.file? and $max.nil?

Thread.new{
  last_count = nil
  while true
    if $count != last_count
      if $max
        STDERR.print "\r#{$count}/#{$max} [#{$count*100/$max}%]"
      else
        STDERR.print "\r#{$count}"
      end
      last_count = $count
    end
    sleep 1
  end  
}

begin
  while data = ($bytes ? STDIN.read(2**12) : STDIN.gets)
    STDOUT.print(data)
    $count += $bytes ? data.length : 1
  end
  STDERR.print "\n"
rescue Errno::EPIPE
end

Explanation time:
  • You can use this script at any point of the pipeline. foo | bar | progress | blah.
  • The script will take advantage of the fact that its STDERR is still linked with your terminal, and output progress information there. It will then clear and overwrite the same line every second with new information.
  • There is no support for multiple progressbars in one pipeline - his is not terribly difficult to do (one would be master, others would send info to it via socket based on tty's inode), but I never found a good use case for it, so I never bothered implementing it.
  • progress script works in two modes - by default it counts bytes (-b), but it can count lines (-l) as well.
  • You can specify what counts as 100% if you want percentage information - with progress -l 1234, progress -b 700m etc. If you specify wrong size of course you get garbage.
  • If you operate in the default byte mode and input is a file - the script will figure out file size automatically. This doesn't happen in line mode, as it would require a potentially expensive wc -l - it's easy to do it manually if you want.
Here's a "screenshot":
$ ./progress < kubuntu-10.04-beta2-desktop-amd64.iso | openssl md5
287244288/708704256 [40%]

Enjoy.