Showing posts with label original research. Show all posts
Showing posts with label original research. Show all posts

Sunday, May 20, 2012

Phosphorus, detergent, and Canada's Experimental Lakes

ResearchBlogging.org
I'm angry at the people who decided that phosphate was growing algae. I'm not sure that I believe that.  –Sue Wright, Texas
Sue Wright, quoted above, was upset because in 2010, sixteen American states banned the sale of dishwashing detergent containing high levels of phosphorus, an aquatic pollutant that sometimes causes eutrophication (algal blooms). Unfortunately, phosphorus is a rather effective component of detergent, so phosphorus-free dishwashing detergents did not immediately perform quite as well as their predecessors. This led some consumers (like our pal Sue) to complain to detergent manufacturers, state governments, consumer protection agencies, and the media.

What I like most about Sue’s complaint is that her anger was directed toward “the people who decided that phosphate was growing algae” rather than the policymakers who drafted and enacted the legislation. Her implied logic is exquisite – a factual claim has resulted in legislation that negatively affects some aspect of my life, therefore I don’t believe this factual claim and furthermore am angry at those who made it!

So, who specifically should Sue have directed her anger toward? Which jackass scientist “decided that phosphate was growing algae”?

The answer, unsurprisingly, is that many independent studies (involving various research groups) have demonstrated that phosphorus pollution, under some conditions, will stimulate algal growth and lead to eutrophication (see Schindler 2006 for a review). Here, I will focus on just one of these studies, perhaps the most influential.

My real motivation for discussing this particular paper is the recent announcement that the Canadian Government is discontinuing its operation of the Experimental Lakes Area (ELA), a collection of 58 pristine lakes that for over 40 years have been set aside for long-term ecosystem monitoring and ecosystem-scale experiments (more on the ELA later).

Green sludge

In the 1960s and 70s, many North American rivers and lakes, especially the Great Lakes, were experiencing rapid declines in water quality (see here and here). Industrial and municipal effluents were stimulating the growth of algae and other aquatic plants (termed ‘eutrophication’) leading to unsightly mats of green sludge, oxygen depletion, massive die-offs of fish and other aquatic life, and problems with the taste and odour of municipal drinking water.

The August 1969 issue of Time Magazine describes the then deteriorating state of Lake Erie:
Each day, Detroit, Cleveland and 120 other municipalities fill Erie with 1.5 billion gallons of inadequately treated wastes, including nitrates and phosphates. These chemicals act as fertilizer for growths of algae that suck oxygen from the lower depths and rise to the surface as odoriferous green scum. Commercial and game fish … have nearly vanished ... Weeds proliferate, turning water frontage into swamp. In short, Lake Erie is in danger of dying by suffocation.
The public, industry, and all levels of government agreed that something had to be done to curb the declining state of North American waterways. However, there was disagreement over the most effective course of regulatory action because at the time, scientists and policymakers were still debating which nutrients were responsible for eutrophication. Was algal growth primarily limited by carbon, nitrogen, or phosphorus?

Schindler 1974

Experiments are the best way to establish causation, but are not always feasible. For example, the best way to test the anthropogenic climate change hypothesis would be to release copious quantities of greenhouse gas into the atmospheres of a random sample of earth-like planets, leave another randomly-chosen bunch of planets untouched, and then compare change in climate across the two groups of planets. Clearly this is not feasible, and clearly we can’t experimentally pollute a bunch of lakes just for the sake of science. Right? Wrong. Well, wrong to the second assertion at least.

The aforementioned Experimental Lakes Area is (was) a wonderful place where scientists could manipulate whole lakes to test hypotheses on the scale of entire ecosystems. In the late 1960s and early 70s, David Schindler – a Canadian limnologist who at the time was director of the ELA – oversaw a number of whole-lake experiments designed to determine which nutrient (out of carbon, nitrogen, and phosphorus) was primarily responsible for eutrophication.

In an initial experiment, Schindler et al. added copious amounts of nitrogen and phosphorus to Lake 227 which naturally had an extremely low concentration of dissolved carbon. If algal growth was primarily limited by carbon (and not nitrogen or phosphorus), then the N + P treatment should not stimulate the growth of algae. However, this was not the case. Within weeks of the treatment, Schindler et al. observed that Lake 227 “was transformed into a teeming, green soup” with algal concentrations up to two orders of magnitude higher than nearby untreated lakes. Clearly, low levels of carbon had not been limiting the growth of algae.

In a second experiment, Schindler et al. divided another lake, Lake 226, into two equal halves using a large vinyl curtain that was sealed into the sediment and surrounding bedrock. The team added an equivalent amount of carbon and nitrogen to both halves of the lake, but added phosphorus to only one side. This manipulation resulted in what James Elser at Arizona State University has called “the single most powerful image in the history of limnology”.

Figure 1. Lake 226 following fertilization with carbon, nitrogen, and phosphorus (below divider) versus carbon and nitrogen only (above divider).

Just a few months after the nutrient additions began, the side of the lake receiving C + N + P was completely covered by a bloom of blue-green algae whereas algae levels on the C + N side were essentially unchanged from when the nutrient additions began. It was abundantly clear that phosphorus had been limiting the growth of algae in Lake 226.

In a final experiment, Schindler et al. manipulated a third lake, Lake 304, to test whether, and how quickly, a lake could recover from phosphorus-induced eutrophication. The team measured the concentration of algae in Lake 304 at approximately monthly intervals over the course of five years, between 1969 and 1973. For three of those years, 1971–1973, the lake received additions of carbon and nitrogen, and for two years, 1971–1972, also received phosphorus. The experiment therefore mimicked what might happen if governments took steps to limit the amount of phosphorus entering a polluted water body. The general finding was that summertime algal concentrations increased dramatically in 1971 and 1972 when the lake was being fertilized with C + N + P, but returned to near baseline levels in 1973 after phosphorus fertilization was discontinued.


Figure 2. Chlorophyll a concentrations (a proxy for algal growth) in Lake 304 from 19691973. Boxplots are based on data extracted from Figure 2 of Schindler 1974 and only include samples from mid-summer; between June and September.

This result again demonstrated that algal growth was limited by phosphorus, and furthermore showed that reducing the amount of phosphorus entering a polluted lake could lead to rapid recovery.

This series of experiments led by Schindler was instrumental in convincing scientists, governments, and the public that phosphorus played a significant role in eutrophication and should therefore be regulated. Throughout the 1970s and 80s, the Canadian government and many American states enacted legislation banning or limiting the use of phosphorus in laundry detergents and other cleaning products.

The Experimental Lakes Area

Schindler’s work on eutrophication represents just a small fraction of the world-class research conducted at the Experimental Lakes Area. Over the past 40 years, research carried out at the ELA has led to 676 peer-reviewed publications including 8 papers in the journal Nature and 15 in Science (the most prestigious scientific journals). In addition, 116 graduate theses and 158 technical reports have been based on research at the ELA.

Much of the research carried out at the ELA has been highly relevant to public policy. Scientists with the Department of Fisheries and Oceans and researchers from universities across Canada and the world have used the ELA to determine how aquatic ecosystems are impacted by things like synthetic hormones from birth control pills, acid rain, aquaculture, common forms of habitat destruction such as the removal of aquatic vegetation, hydroelectric reservoirs, eutrophication, and the accumulation of heavy metals and organic toxicants. A news article in the journal Science refers to the ELA as “Ecology’s supercollider” and James Elser of Arizona State University suggests that “it’s hard to overstate the impact [the ELA] has had”.

Tragically, the Canadian government feels that $600,000 per year (the ELA’s estimated operating budget) is too high a price to pay for world-class environmental science. In case it isn’t clear, I emphatically disagree.

____________________________________

Schindler, D. (1974). Eutrophication and Recovery in Experimental Lakes: Implications for Lake Management Science, 184 (4139), 897-899 DOI: 10.1126/science.184.4139.897

Tuesday, December 13, 2011

Facial metrics predict unethical behaviour

ResearchBlogging.org
At the sight of that skull, I seemed to see all of a sudden ... the nature of the criminal – an atavistic being who reproduces in his person the ferocious instincts of primitive humanity and the inferior animals. Thus were explained anatomically the enormous jaws, high cheek-bones, prominent superciliary arches ... found in criminals, savages, and apes ... and the irresistible craving for evil for its own sake.   
–Cesare Lombroso
The idea that facial metrics can provide information about a person’s propensity for unethical behaviour is making something of a comeback in the scientific literature. Recent research has correlated facial width-to-height ratios with:
(a) time spent in the penalty box among Canadian hockey players (Carré and McCormick 2008),
(b) propensity to punish an opponent at a cost to oneself (Carré & McCormick 2008), and
(c) the propensity to selfishly exploit an opponent’s trust in a game context (Stirrat & Perrett 2010).

Not only can these “unethical” or “aggressive” tendencies be correlated to facial metrics via meticulous research – it turns out that casual observers are also able to predict a complete stranger’s propensity for aggression with surprising accuracy based only on a picture of that person’s face (Carré and McCormick 2009). The ethical implications of these findings are interesting and perhaps controversial (think racial and criminal profiling), but I will resist the temptation to go down that road at the moment, and instead focus on a recent study by Michael Haselhuhn and Elaine Wong titled Bad to the bone: facial structure predicts unethical behaviour.

Three things in particular attracted me to this paper: (i) the author’s use of the term ‘unethical’ which in my mind is subjective and imprecise; (ii) their reference to William March’s classic novel The Bad Seed and its conclusion that some people are just ‘born evil’; and (iii) the final sentence of the discussion “Perhaps some men truly are bad to the bone.” In truth, it was ridiculousness that drew me in, but nonetheless, let us proceed with the science!

Haselhuhn and Wong summarize their main result as follows:
...we show that genetically determined physical traits can serve as reliable predictors of unethical behaviour ... Specifically, we identify a key physical attribute, the facial width-to-height ratio, which predicts unethical behaviour in men.
In a first experiment, the authors arranged for 96 pairs of Business Admin students to participate in a negotiation exercise (conducted via email). Each pair consisted of a ‘buyer’ and a ‘seller’. Sellers were told that the property they were selling must not be commercially developed whereas buyers were instructed to obtain the property specifically for commercial development. The researchers quantified ‘unethical behaviour’ based on whether or not a buyer explicitly misstated his or her intentions at any point during the email negotiations.

In a second study, 103 undergraduates completed an online survey designed to assess their own ‘sense of power’ and then were allowed to enter a lottery for a chance to win a $50 gift card. The number of times that a student could enter the lottery was to be determined by a single roll of two dice (simulated at random.org). Because there was no oversight when students inputted the result of their dice roll at the end of the survey, there was nothing stopping participants from cheating and entering a higher number than was actually rolled (thereby increasing their chance of winning the gift card). Furthermore, because the dice roll was random, researchers were able to quantify cheating (at the group level, but not the individual level) based on deviations from the expected average roll of 7. For both experiments, two research assistants measured facial width-to-height ratios (hereafter ‘facial WHR’) of all participants based on school photographs (inter-rater agreement was high, r = 0.758).

So, what happened? In the first study, 18/96 buyers engaged in explicit deception during the email negotiation. Based on logistic regression, the probability of deception significantly increased with increasing facial WHR for men, while women’s facial WHR was not significantly related to deception. In the second study, the average reported dice roll from 103 participants was 7.76, significantly greater than what would be expected by random chance (i.e. 7.00), and therefore indicative of cheating. The reported dice roll (and therefore the incidence of cheating) did not differ between men and women. Just as before, ‘unethical behaviour’ (reported dice roll) was significantly and positively related to facial WHR for men, but not significantly influenced by facial WHR for women.

Recall that in the second study, participants also completed a survey about their own sense of power. It turns out that sense of power was positively related to both facial WHR and reported dice roll for men but unrelated to both variables among women. Some statistical voodoo (similar to path analysis) allowed the researchers to conclude that sense of power mediated the relationship between facial WHR and cheating behaviour among men.

Holy eff, right? I’m pretty amazed by these data. Before reading the paper, I expected that I would take issue with their analyses or conclusions, but the study was in fact well done. The only issue I take is with the authors’ claim (referring to the second experiment) that
...our approach introduces potential noise to the data as, for example, men with smaller facial WHRs may legitimately roll and report higher dice totals. Thus, testing for cheating behaviour using this paradigm represents a conservative test of our hypotheses.
This is simply untrue. It is equally likely that men with larger facial WHRs might legitimately roll and report higher dice totals. Their approach is imprecise, but not ‘conservative’. A more precise way to test the relationship between facial WHR and cheating would be to use the difference between reported dice rolls and actual dice rolls as a dependent variable. This would require an experimental modification allowing the researchers to know with certainty what each participant actually rolled (e.g. using cameras, software tracking, etc.). Nonetheless, the results seem robust.

So then, how could this relationship have evolved? Intuition would suggest that physical signals that reliably predict unethical behaviour should be selected against. If you’re trying to deceive or cheat someone, you don’t want to tell them up front. Haselhuhn and Wong suggest that a relationship between certain facial features and unethical behaviour could evolve through pleiotropic associations with sexually selected traits such as aggression and dominance. If females like to mate with dominant males, and dominance is correlated both with certain facial features and unethical behaviour, than unethical behaviour may come to be associated with those same facial features.

Importantly, the correlation between facial WHR and unethical behaviour demonstrated by Haselhuhn and Wong does not necessarily imply that some people are ‘born evil’. For starters the effect sizes they reported were fairly small – facial WHR explained a relatively small proportion of the variance in propensity to cheat and deceive. Second, as the authors point out
...it is important to recognize that other developmental processes may play an important part in forming these links ... one possibility is that men with greater facial WHRs are perceived and treated by others in ways that encourage unethical action (i.e. a self-fulfilling prophecy).
Again, although the reported relationships seem to be robust, I see no data here or elsewhere supporting the idea that some men are “bad to the bone”.

____________________________________

Haselhuhn, M., & Wong, E. (2011). Bad to the bone: facial structure predicts unethical behaviour Proceedings of the Royal Society B: Biological Sciences DOI: 10.1098/rspb.2011.1193

Monday, December 12, 2011

Too sexy for my smile

CBC News recently reported on a study published in the journal Emotion that attempted to determine how body language influences perceived sexual attractiveness. I take issue with some of the methods and interpretations. You can read the actual paper here, and the CBC article here, but the gist of it is as follows:
  • A large sample of men and women were shown photographs of members of the opposite sex and asked to rate their sexual attractiveness.
  • Each photo depicted a ‘model’ displaying one of four emotions – happiness, pride, shame, or neutral.
  • In the first study, all participants were asked to rate a single photograph. All male participants rated the same female model in one of the four possible poses. Likewise, all female participants rated the same male model in one of the four poses.
  • In a second study, three large groups of participants rated a bunch of photographs that were viewed online. Again the photographs displayed a member of the opposite sex expressing one of the four emotions. This time, however, the photographs (over 400 of them) were obtained online (e.g. from Google Images) and sorted into their respective categories (2 genders • 4 emotions = 8 categories) by trained assistants according to published guidelines. So, unlike the first study, each category here contained pictures of many models, and different models were used to depict each emotion.
  • The general result that held across both studies was that males expressing happiness were rated the least attractive and males expressing pride were the most attractive. The trend was essentially reversed for female models, such that happy females were rated the most attractive whereas females expressing pride were among the least attractive.
  • There were other interesting results and many details I have left out for the sake of brevity. Check out the original paper for more information.

So what are the shortcomings of this study? My problem with the first study (which in fairness the authors do acknowledge) is that the sample size for each gender is one. All female participants rated the same male model, and all male participants rated the same female model. This study provides great evidence that this particular woman and this particular man are respectively more and less attractive when smiling, but we have no evidence that this trend exists in the population at large. It is entirely plausible that for different subjects the trend would be reversed.

To really hammer this point home, consider the question – are songs in the key of C minor more enjoyable than those played in the key of D minor? What the authors have essentially done is asked the London Philharmonic to record two versions of Beethoven’s Symphony No. 5 – one version in the original C minor, and the other transposed into D minor. They then asked 184 participants to rate the enjoyability of one of the versions, and concluded that songs in C minor are more enjoyable than songs in D minor because participants on average gave the C minor version of Beethoven’s Symphony No. 5 a higher enjoyability score. Crazy, right!? Maybe Beethoven’s other symphonies actually sound better in D minor, or maybe his symphonies sound better in C minor but his sonatas are more enjoyable in D minor, or maybe Beethoven’s compositions are generally more enjoyable in C but Bach’s are consistently more enjoyable in D, etc. Point is, you can’t make generalizations with a sample size of one. Again, the authors do actually acknowledge this point, and claim that the second study addresses this shortcoming.

Problem number deux. In the second study, where photographs were obtained from the internet and many models were used in each category, I believe there were systematic differences between categories apart from just emotional expression. Admirably, the authors have posted all of the photos used in their study here. There are a few trends that really stuck out for me. One is that photographs in the pride samples were mostly comprised of athletes in their race or match apparel, whereas few or no athletes appeared in the other three categories. Another trend is that most neutral photographs tended to show only the face and sometimes shoulders, whereas hands, upper bodies, and even full bodies appeared in the other categories. A third issue is that neutral faces were almost always facing directly toward the camera with no angle or tilt, whereas faces and bodies in other categories were much more likely to be angled. There also seem to be differing proportions of professional-looking photographs between the different categories (the authors did partially control for the number of models that appeared to be professional models, but only in two of the three samples). In sample A, all of the shame photographs appear to be professionally taken, whereas most of the neutral photographs appear to have been taken by a kid at the DMV.

Going back to the music analogy, the authors have essentially downloaded a bunch of songs from iTunes, half in C minor and half in D minor, but for whatever reason most of their C minor songs happen to fall into the Hip-Hop & Reggae genre, and most of the songs in D minor happen to belong to the Country & Western genre. Even if we have a large and random sample of the population rating the enjoyability of these different songs, any average differences observed between songs in C and D minor are not necessarily due to the different key signatures, but could just as easily be due to any of the myriad differences that (on average) distinguish Hip-Hop music from Country music. Of course the same is true for the different sets of photographs in the study described above, except the confounding variables in this case were photograph quality, angle of head from camera, proportion of body appearing in the photograph, clothing and location of the model, etc.

To conclude (finally!), I don’t really doubt the claims made in this study, I just don’t think they necessarily follow from the obtained results. There are logistical limitations to any study, and we can rarely design studies that will definitively test a hypothesis of interest while controlling for every possible confounding factor. I do however think that it is reasonable and possible to more conclusively and meticulously test the hypothesis that emotional expression influences perceived attractiveness by members of the opposite sex.

____________________________________

Tracy, J. L., & Beall, A. T. (2011). Happy guys finish last: the impact of emotion expressions on sexual attraction. Emotion 11:1379-1387. DOI: 10.1037/a0022902