Monday, July 7, 2014

RCTs are Necessary to Determine Causality. Right? Wrong.

It seems to be a widely held belief that randomized control trials are preferable if one is interested in determining causality.  Causality is something of a slippery concept, but let us use the following definition:  Drug A causes an increase in survival if drug A increases survival for at least ONE person relative to drug B.

I think this is a minimal requirement.  If it is not true that drug A increases survival for at least one person relative to drug B, then drug A certainly does not cause an increase in survival.

If we observe the proportion of patients who survive on drug A and the proportion of patients who survive on drug B, can we determine if drug A causes an increase in survival?  

No.  We don't have enough information.

If more patients who take drug A live longer than patients who take drug B, we cannot (without more information) determine why they lived longer and whether it had anything to do with drug A.

What is the minimum amount of information we would need to determine if drug A causes greater survival?  The answer is somewhat technical, but there are two cases where we could make the determination. 

The first case is where we observe patient survival for a representative group of patients who only had drug B available.  For example, prior to 2004 stage III colon cancer patients received 5FU as adjuvant therapy as oxaliplatin had yet to be approved by the FDA.

The second case is where we are willing to assume that patients are assigned to the drug for which their survival probability is higher.  This is what is called a "behavioral" assumption.  Such assumptions are relatively common in economics, but generally frowned upon outside of economics.

In both of these cases it is possible to determine whether drug A causes greater survival relative to drug B by simply looking at the probability of survival between those patients that take drug A and those patients that take drug B.

It is not necessary to have some ideal randomized control trial if we are able to observe survival for a group of representative patients who had no access to drug A or if we are willing to assume that patients are not assigned to the drug that is more likely to kill them.


Tuesday, June 17, 2014

Solving the Wrong Problem


For the last 100 years or so statisticians and econometricians have spent all of their energy solving the wrong statistical problem.

In statistics we are interested in determining what happens to some outcome of interest (Y) after making a change to some other observable variable (X).  For example, we are interested in increasing survival from colon cancer (Y to Y') using some new drug treatment (X to X').  The problem is there is some unobservable characteristic of the patient or the drug or both (U) that may determine both patient survival and the use of the new drug treatment.

There are two statistical problems.  

The first one, the one we spend all our energy on, is called "confounding."  In the picture above this problem is represented by the line running from U to X.  The unobserved characteristic of the patient is determining the treatment the patient receives.  In this paper, there seems to be a tendency for oxaliplatin to be given to healthier patients which may explain the survival difference between the oxaliplatin group and the non-oxaliplatin group.  To solve this problem we spend millions and millions of dollars every year to run randomized control trials.  In economics, we devise fancy and clever ways of overcoming the confounding with instrumental variables.

The second problem doesn't really have a name.  I will call it "mediating."  This problem is represented in the picture by the line from U to Y.  The problem is that there may be some unobserved characteristic of the patient that is mediating the effect of the treatment on the patient's survival.  Oxaliplatin may have greater effect on survival for younger patients relative to older patients (see here).  We do not spend any money or much time devising ways to solve this problem.  In fact, we often give up before we start by saying that it is impossible because it is not possible to observe the same patient's outcome under two different treatments.


The problem with spending all our time on the first problem is that once we solve it we are still no closer to solving the second problem and we still don't know what will happen to any patient when given the treatment being studied.

In the graph to the left, the outcome (Y) is a function of both the treatment (X) and the unobserved patient characteristic (U).  Although the experiment removes the line between U and X, the line between U and Y remains.

We can conduct as many experiments as we like and still be no closer to knowing what will happen when we give the treatment to a new patient because we don't know anything about that patient's unobserved characteristics.

Saturday, June 14, 2014

I Love Polynomials

The statistician, Andrew Gelman, hates polynomials.  I love them.

Polynomials are really very cool and they have a lot of very nice properties.

One of the most important properties they have in statistics is that they form a "ring".  This is an algebraic term meaning that any two elements of the set may be added (subtracted) or multiplied (divided) and the corresponding outcome is also an element of the set.  Fractions form a ring.  If you take any fraction and multiply it by any other fraction you get a fraction (a rational number).  Numbers (natural numbers) do not form a ring.  5/4 is not a natural number.

What is the big deal about rings?

The big deal is that if we have a set of continuous functions on a closed and bounded space that form a ring, then that set can approximate any (that is any and all) continuous functions on the aforementioned space.  In English.   If we want to approximate a continuous function, then we can't do any better than to use a polynomial.  For those interested, check out Stone-Weierstrass Theorem.

Polynomials are also unbelievably easy to estimate.  We can just do ordinary least squares regression and viola.

So any function you want to estimate can be approximated with a polynomial and they are really easy to estimate.  What is not to love?

How do you know that the polynomial you are estimating is really equal to the function you are interested in approximating?  Well.  You don't.

Gelman points out that there is a tendency to estimate very high order polynomials and he discusses the implications in this unpublished working paper.  The problem is that the data may not allow such polynomials to be identified.  The result is that many important coefficient estimates are simply made up numbers.  

I show in this paper that if you have a sequence of polynomials and a data set with enough information to accurately each polynomial in the sequence, then that sequence converges to the function of interest.  I also show in a Monte Carlo (with made up numbers) experiment that if there is not enough information in the data to accurately estimate a high order polynomial, the approximation error is very large.

The corollary is that if you have a polynomial that is a poor approximation of the function of interest then there is not enough information in the data to accurately estimate the high order polynomial.

Sunday, June 8, 2014

RCTs are Like Looking for Money Under the Street Light

Mutt and Jeff (June 3 1942)
Rubin (1974) argues that we should prefer randomized control trials (RCT) to observational data because the "casual effect" of the treatment is measured from the RCT data under mild assumptions. 

Rubin defines the "causal effect" of a new treatment as the difference between the outcome of the patient when she is given the new treatment and the outcome of the patient if she had received the current standard of care.  Rubin acknowledges that the difference is not observed and is in fact unobservable.  A patient can only ever receive one treatment and so it is not possible to observe the outcome in two alternative treatments.

Like the man in the top hat, Rubin suggests looking for the information in the light.  In Rubin's case the "light" is provided by the RCT which measures the average treatment effect under mild assumptions.  Rubin argues that average treatment effect is a measure of the difference in outcomes for the "typical" patient.  If we take "typical" to mean that it is true for some reasonable sized group of patients, then there is no reason to believe that the "typical treatment effect" will even have the same sign as the average treatment effect.  The average treatment effect averages over the difference in outcomes for each of the patients.  If some patients benefit from the treatment and some patients are harmed by the treatment then the average treatment effect may be positive or negative depending on the relative sizes of the two patient groups and the relative sizes of the benefits or harms.

If the average treatment effect is positive then we know for certain that there exists one patient for whom the new treatment was better than the existing treatment.  That is it.  That is all we know for certain.  It may be that all patients are better off with new treatment or it may be that (almost) all patients are worse off with the new treatment.  The average treatment effect is observed by the light of the RCT but it tells us very little about what we are looking for.

Saturday, May 31, 2014

If Wishers Were Horses, Beggars Would Ride

The National Cancer Institute (of the NIH) announced this week a large reorganization of its clinical trial system.  As part of the reorganization it announced smaller budgets for running clinical trials.  This reorganization has been coming down the pike for a while now and the smaller budgets are a matter of fact given reduced funding from Washington.

What I found disturbing in the announcement was the repeated claim that new technologies in oncology drugs would reduce the need for large clinical trials.  

According to the announcement

Although the screening tests may need to be performed on very large numbers of patients to find those whose tumors exhibit the appropriate molecular profile, the numbers of patients required for interventional studies are likely to be smaller than what was required in previous trials. That is because the patient selection is based on having the target for the new therapy, leading to larger differences in clinical benefit (such as how long patients live overall or live without tumor progression) between the intervention and control groups.

It is true that breakthrough advances such as the AIDS cocktail or Gleevac can show themselves to be enormously effective even in small trials, but that doesn't mean that we should expect all new drugs or treatments coming into development to be breakthroughs.

It is unclear to me why we should expect that targeted therapies should have larger effects on survival.  I see why we should expect targeted therapies to be more targeted and thus only likely to work for a small subset of patients with very particular genetic mutations in their tumor.  But even if the therapy works on the "bench" it may not have the same effect once it is put into humans.  As the announcement states, targeted therapies will require greater amounts of genetic screening in order to find the right patients.  More over, the total population of patients with a particular genetic mutation may be extremely small.  The future of targeted therapies may well involve smaller clinical trials, but I think NCI is being rather optimistic believing that we won't need large trials.  Today's "Daily News" from ASCO presents a very different view from Don Berry.

Gleevac is the poster boy for new age of targeted therapy, but it is a drug that seems to be exception rather than the rule.  Gleevac was able to solve a very particular genetic problem for a very particular class of cancer patients.  The genetic problems in most common cancers seem to be substantially more complicated and do not seem amendable to single target therapies.  In colon cancer, genetic testing is required for certain drugs, not because these drugs have amazing breakthrough effects, but rather because they don't seem to work when certain genetic mutations are present (see here).

In the mean time non-targeted therapies like immunological therapies are starting to be developed.  Will these therapies also require smaller trial sizes?

Let's hope NCI gets its wishes and all future drugs are breakthrough therapies that don't require large clinical trials and beggars can finally ride.

Wednesday, May 21, 2014

A New Way to Solve Confounding?

IV Graph from Imbens (2014)
Confounding refers to statistical problem that there is some unobserved characteristic of the patient that is both determining the patient's observed treatment and the patient's outcome.

For example, this study shows that older stage III colon cancer patients are much less likely to receive oxaliplatin as an adjuvant therapy than younger patients.  This may be the reason that in the Medicare data, oxaliplatin is associated with bigger survival effects than in the randomized control trials.  The Medicare data suffers from a confounding problem.  Doctors of sicker patients may be less willing to prescribe oxaliplatin because of its side effect profile.  The observed difference in survival may not be due to the use of oxaliplatin, it may simply be the fact that the non-oxaliplatin patients are sicker.

In the graph to the right, the unobserved variable (patient "sickness") is represented by the red U.  The patient's treatment (oxaliplatin or not) is represented by the black X and the patient's survival is represented by the black Y.  We would like to know whether there is a blue line from X to Y, representing treatment effect of using oxaliplatin on survival.  But we can't determine the treatment effect because U is affecting both X and Y through the red lines from U to X and U to Y.  Sicker patients are less likely to get oxaliplatin (red line from U to X) and sicker patients have lower survival (red line from U to Y).

A standard way to solve the confounding problem is to observe (or introduce) a fourth variable (Z) which is called an "instrumental variable."  As the graph shows, the instrument is some observed characteristic of the patient that determines the patient's treatment choice but is unrelated to the patient's unobserved characteristic or the patient's survival.  In randomized control trials the instrument is the random number generating process that is used to assign patients to treatment arms.

In the Medicare data on the use of oxaliplatin, the instrument may be the date of the diagnosis.  Patient's diagnosed earlier were much less likely to receive oxaliplatin than patient diagnosed at a later date.  By looking at changes in survival over the time period of the introduction of oxaliplatin we can determine the causal effect of oxaliplatin on survival (assuming no other major changes to treatment during the same time period).  

An alternative way to solve the confounding problem is to measure all the confounding characteristics.  If we observe U then we can simply measure the effect of X and U on Y.  If we observe the co-morbidities of the patient we can measure the relationship between the co-morbidities and the use of oxaliplatin on survival.  The problem with this approach is that we may not observe all the confounding factors.

A new paper of mine (see discussion here) suggest an alternative approach.  Instead of attempting to directly measure U, we infer U from observable characteristics of the patient.  Instead of attempting to directly measure the "sickness" of the patient, we look at observable characteristics of the patient like their age and use those signals to determine the distribution of patient's latent sickness type.

This mixture model approach has the advantage of not requiring instruments and not requiring that observe every possible characteristic of the patient that may be determining the treatment choice.

Saturday, May 17, 2014

Can Mixture Models Cure Cancer?

Mixture model example from Wikipedia
The other day I was emailing with a friend about my new paper in which I use something called a "mixture model" to analyze the variation in treatment effects.  I'm pretty excited about the idea of using mixture models in this way, but was my friend was somewhat dismissive.

The idea of a mixture model is pretty straight forward.  Imagine observing a distribution where the four distributions pictured to the right were all equally likely and we saw the probability of an outcome less than -2.  For the purple type the probability is 0.5, for the red type it is 0.  For the green type it is about 0.05 and for the blue type it is about 0.1.  So in our data we should see an outcome less than -2 about (0.5 + 0 + 0.04 + 0.1)/4 = 0.16 (about sixteen percent of the time).  

The statistical question of interest is if we observe outcomes less than -2 sixteen percent of the time, can we use this information to decompose the observed probability into the four underlying types that generate the observed data?  Can we tell from our data how many hidden types there are?  Can we tell the proportion of each type?  Can we determine the distribution for each type?

The answer is no.

It is easy to see this.  Imagine that instead of the purple type having a probability of 0.5 that the outcome is less than -2, it is 0.4 and the probability for the blue type is instead 0.2.  The observed probability in the data will again be 0.16.   We can arbitrarily change the underlying type probabilities, as long as the aggregate is 0.64.  All such possibilities are consistent with what we observe.  Similarly, we can change the weights and the probabilities in many different ways and still get the observed sixteen percent we see in the data.

OK, so if we can't decompose our observed data into the underlying distributions, what is the point?

The interesting thing about these models is that with certain information and certain assumptions about how the data is generated, it is possible to decompose the data into the underlying distributions.

Unfortunately, it is very very common to estimate these models when there is not enough data to do the decomposition.  Often the resulting decomposition is coming from arbitrary and non-credible assumptions made by the researcher rather than any actual information in the data.  Worse, it is often unclear how much of what we know about the distribution is due to information in the data and how much is due to the arbitrary assumptions of the researcher.

In 1977, a mathematical statistician, Joseph Kruskal, worked out in this paper, sufficient conditions for the data to provide enough information for the observed distribution to be decomposed into the underlying distributions.  That is, Kruskal presented a set of conditions for when the data and not arbitrary assumptions of the researcher would provide enough information for the decomposition.  More recently, in this paper, signal engineer, Nikos Sidiropoulos, and co-authors presented necessary conditions on the data for the decomposition to possible.

My new paper thinks of their being different types of people, where not only may these different people have different outcomes, but the treatment being tested may have different effects.  When we test a new drug using randomized control trials we generally present the results aggregated over the different types.  If we find that the drug increases survival, we do not know if it increases survival for some people, all people, or most people.  The objective of the statistical analysis is to uncover the different hidden types and ultimately to target particular drugs to particular sub-groups of the population.  The hope is that this can be done without resorting to arbitrary and non-credible assumptions.  My friend remains skeptical.