Showing posts with label Bowel_Cancer. Show all posts
Showing posts with label Bowel_Cancer. Show all posts

Thursday, 20 October 2011

Sunlight and bowel cancer incidence

I mentioned a fortnight ago that I was going to investigate a possible relationship between the regional distribution of bowel cancer incidence in the UK and of sunlight.  And so I have.

I've chosen to work with incidence rather than mortality, because the incidence numbers are larger and hence less affected by random variation, and because it seems to me that sunlight is more likely to affect incidence than outcome.

The hypothesis is that sunlight is linked to bowel cancer incidence because exposure to ultraviolet light promotes synthesis of vitamin D, which has a protective effect.  It's the UVB (higher energy UV) component of the light that's important for vitamin D synthesis, and the process is not linear.  However, for a first pass at least, I've used data I could quite easily get online, which is sunlight incidence from this site, intended for use in solar energy calculations (I told it I'd use horizontal panels).  For each district I queried for one or more population centres, averaging the results if necessary.  The solar incidences I found varied from 2.14kWhr/m^2, averaged over the whole year, in Shetland, to  3.00 in Kingsbridge, Devon.


Here's a scatter plot of the data.  I've used a non-zero intercept on the sunlight axis.  I've resisted the temptation to let Excel draw a regression line.

It looks as if there might be something there.  I calculated a population-weighted correlation: the result is -0.300 .  To check for significance I tried shuffling the sunlight data so that the numbers were applied at random to the districts, recalculating the correlation each time.  In 1000 tries none of the correlations where anywhere near that large: there may be a lot of perturbations that would give that much correlation, but they're a tiny fraction of the 380! available.

However, this is an in-sample test - I've used the same data to form the hypothesis and to test it.  To avoid that, I tried ignoring the eleven labelled points I looked at when I came up with the notion.  The population-weight correlation was reduced to -0.256 .  This is still larger than anything I got in 1000 shuffles of the data (the largest therein being -0.199).  I conclude that there is a genuine negative correlation in the UK between bowel cancer incidence and sunlight.

Correlation does not prove causation - I suppose I would have got a similar result if I'd looked at latitude instead of sunlight, and one could attribute that to lifestyle differences between different parts of the country.  , But more rigorous analyses in the USA and Japan have found a negative correlation of bowel-cancer mortality with sunlight even after correcting for dietary differences.

Population-weighted linear regression gives a regression slope of -15.34, i.e. the best guess is that an increase in solar incidence of 1kWhr/m^2 reduces bowel cancer incidence by about 15 per 100,000 people.  To give an indication of the potential effect, I worked out the effect of applying an adjustment in each district as if every district had the solar incidence of the sunniest.  That would reduce bowel cancer incidence by 4298 cases per year.  If the ratio of mortality to incidence were maintained (research suggests sunlight might in fact improve survival) that would save 1742 lives per year.

I am not proposing that we promote climate change to allow more sunlight through the atmosphere.  But it might be possible to encourage people to expose themselves to sunlight more (not of course neglecting its dangers: there were 2067 deaths from malignant melanoma in the UK in 2008), or dietary vitamin D might have a similar effect (the data are suggestive but unclear).

Wednesday, 5 October 2011

Funnel plot of UK bowel cancer incidence

If you've already read more than you want to about colorectal cancer, please look away now.
This chart is a funnel plot of incidence (i.e. new cases) rather than mortality.  It shows that there is significant variation from one district to another - there are several points outside the dashed lines.

It would be a mistake to deduce from this and the previous analysis that incidence depends on where you are but mortality doesn't.  We didn't prove the null hypothesis for mortality, we just failed to disprove it (outside Glasgow).  But it's easier to get statistically significant results for incidence, because the numbers are larger.  Strangely, the best evidence that mortality varies with district may be in the incidence numbers not the mortality numbers.

However, the data are for diagnosed incidence, not for the (unknown) total incidence.  Starting a screening programme will cause a jump in diagnoses, which will mostly but not entirely fall off with time - in the absence of screening, some people with bowel cancer may live out their lives without being diagnosed, and die of something else.  So for example the high incidence of diagnoses in Wirral in 2008 may be more to do with the start of a screening programme in 2007 than with a high total incidence.

I note that all the points well above the upper dashed line lie in the north of the UK, and all but one of the points well below the lower dashed line lie in the south of the UK.  I asked an oncologist, who told me that bowel cancer is known to be negatively correlated with sunlight, more so than any other cancer.  I'll try to show to what extent this factor is explanatory.  Meanwhile, here's a US map, generated on this site, of the sort that was first used to generate the sunlight hypothesis.

Tuesday, 4 October 2011

Age Standardization

I accused Beating Bowel Cancer of doing its age standardization wrong: I withdraw that.  What they seem to have done is to standardize against European norms, without saying so - I realised this when I had another look at the Cancer Research UK spreadsheet, which gives a very similar standardization.

Actual bowel cancer mortality across the UK is substantially higher than the numbers from Beating Bowel Cancer tell us, which are for a hypothetical population with fewer old people in it.

Tuesday, 27 September 2011

Bowel Cancer Statistics - a funnel plot


I'm grateful to David Spiegelhalter of Understanding Uncertainty for the suggestion that I should display these data in a funnel plot.  (Click on the plot for a full-screen version.)

For a given population size, and assuming that the death rate is uniform across each authority, an authority has the same probability of falling outside the funnel lines whatever its size - the funnel is narrower at the right-hand end of the plot where the authority sizes are larger and the standard deviation of the distribution is smaller relative to its mean.

Under the uniform distribution assumption, the probability of a point falling above the upper dotted line is 2.5%, as is the probability of a point falling below the lower dotted line.  There are 378 points, so typically nine or ten of them would be above and below the dotted lines.  For the dashed line, the probabilities are 0.2646%, which I chose so that one point would typically be above and below the lines.

What stands out from the plot is that Glasgow City's poor result is fantastically unlikely to be random.  There may be a meaningful pattern too in the figures for North Lanarkshire and Falkirk, which cover the area from Glasgow east-north-east to the Firth of Forth above Edinburgh.

Only Westminster falls below the lower dashed line.  Since we expect one point there at random, this is perhaps not worthy of much attention.  However, there are a lot of points - 21 - below the 2.5% line, most of them in south-east England (one of them - Stirling - is in Scotland).

I suspect that there are genuine regional variations in outcome, Glasgow aside.  But the data need looking at over regions larger than most local authorities.

(Wales, which is presented as a single, very large, region is a long way off the right-hand end of my plot.  But it falls comfortably within the funnel.)

[This post is a follow-on to my previous analysis of the same data]

Monday, 19 September 2011

Bowel Cancer Statistics - update

I've revised the post below, having done some work on the data.

Thursday, 15 September 2011

'Three-fold variation' in UK bowel cancer death rates

[I rewrote this post on 19th September, having done further analysis of the data.  The overall conclusion is the same, but better supported by the analysis.  I corrected the penultimate paragraph on 4th October.]

I've copied the title from this BBC story.  The story is an uncritical account of a press release by the charity Beating Bowel Cancer.  "Beating Bowel Cancer calculates that over 5,000 lives could be saved every year".

The press release announces an on-line Bowel Cancer Map which allows one to find the (age-standardized) bowel cancer incidence and mortality for each local authority in England and Scotland.  This is based on 2008 figures provided by UKCIS.  (The raw data are available from UKCIS to registered users only.)

The headline finding is that death rates vary from 9.16 deaths per 100,000 in the semi-rural district of Rossendale (in Lancashire) to 31.09 deaths per 100,000 in the city of Glasgow.  I suppose that the calculation that over 5000 lives a year could be saved is based on reducing the death rate nationally from 17.68 to 9.16 - for a population of 61 million that would save 5,197 lives each year (17.68 is my calculation of the UK-wide death rate, using their data.  They give 17.27 for the death rate in England, and higher figures for Scotland, Wales and Northern Ireland).

But this statistical analysis is completely wrong.  Beating Bowel Cancer didn't reply to a request I sent them for a spreadsheet containing the numbers numbers behind the map, so I've scraped them out by semi-automated postcode query.  The total annual deaths in the data I've got is 15,936 compared with 16,259 reported for 2008 by Cancer Research UK, so I think I've been successful enough in extracting the data.  For a given expected death rate, the actual number of deaths in each district will be a random number sampled from a Poisson Distribution with mean equal to the expected number of deaths in that district.  I assumed that the expected death rate is the same throughout the country, and simulated on a spreadsheet numbers of deaths for every district in the UK. After each simulation, I found the district with the lowest actual death rate in that simulation, and the one with the highest.  The result was that on average the district with the lowest death rate has about 7 deaths per 100,000.  The district with the highest death rate has about 32 deaths per 100,000.  So the observed range - 9 to 31 deaths per 100,000 - is in fact slightly (not significantly) smaller than one would expect if the variation were purely random.  There's nothing in that range to suggest that the expectation is any different from one district to another.

What is happening here is that the analysis has been done over districts that are small enough for random variation to swamp any systematic variation.  Taking this to extremes, one might calculate deaths per household, and find that in the best households no one at all died that year of any cause.  If we could duplicate that for all households, we could all live for ever.

It's worth saying a bit more about the population sizes in each area.  The Bowel Cancer Map doesn't give these numbers directly, but it reports numbers of deaths and "age-standardised" deaths per 100,000.
Age-standardisation adjusts rates to take into account how many old or young people are in the population being looked at. When rates are age-standardised, you know that differences between the local authority areas do not simply reflect variations in the age structure of the populations...
From the data given it is simple to calculate the age-standardized population used for each area.  On this basis, Rossendale has 76,000 people and City of Glasgow has 675,000.  It is not surprising that the lowest death rate is in one of the smaller areas - they are the ones with the greatest random variation.  But if the results are simply random it is surprising that the highest death rate is in a large area.  So I do think that the high death rate in City of Glasgow is not purely random - the Cancer Research UK data confirm that death rates are significantly higher in Scotland.

One more thing.  Adding up the populations for each area, calculated from the deaths and death rates, I get a UK population of 89.75 million.  Since the true figure for 2008 was about 61 million, that's rather surprising.  Doing the same calculation on the data for number of cases, which the map also gives, I get a UK population of 83 million.  The Cancer Research UK table is using a population of just over 61 million, and therefore gets a noticeably higher death rate for about the same number of deaths.  [Update: I find that the CRUK table offers age-standardized rates also: they are very similar to the Beating Bowel Cancer rates.  The heading in the table reads "Age-standardised rate (European)...", so it seems that the standardization is to a European-average age distribution.]

I'm not sure whether I should be railing against the use of stupid statistics in a good cause...