This is part 2 of our series on the complexities of the data behind Delphy, our tool for understanding the evolution of pathogens during an outbreak. In part 1, we describe the type of data Delphy generates. If you haven’t already, it’s probably helpful to read that first. As a reminder, the main takeaways are:
In this post, we’ll describe how to launch a Delphy run and what that looks like. Which, by the way, is very different from what you get with other tools for Bayesian phylogenetics. They run from the command line, so you might get some readouts on the screen or there might be a log file you can parse to evaluate progress. But since there’s no insight into how it’s progressing, people tend to set the job for a predetermined number of trees, wait hours or days until it’s done, and then take a look at the results. If you happened to have some bad data, that’s a lot of time wasted. In Delphy, the same screen that launches the run has numerous indicators of how the job is progressing. And you can pause the job at any moment to check the data, and restart it if all looks good.
After sequences are uploaded and parsed in Delphy, you are automatically taken to the “Trees” tab.
On the left side, you get to see a first guess of what the evolution of the pathogen looks like. Delphy shows an initial tree, with the root in the upper left, branching towards the right. Each tip where the branches come to an end represents one of the uploaded sequences. Until the user presses the run button (it looks like a play button), we only have this one tree1.
In the app, to the right of the tree, there are charts for various “traces” generated during the run. Each trace is a particular measure of a tree, and the charts track that measure across all the base trees that Delphy generates. For example, since each base tree is a take on the evolution of the pathogen, we get a mutation rate for each one of those trees. The “Mutation Rate μ” trace chart allows us to see all those rates together and compare them across the entire Delphy run. Most of the traces have two charts: a seismograph-like chart showing the values for each base tree (there’s only one value at the start), and below that, a distribution that summarizes the values. They start with only one value each, but they fill up as the run progresses. All together, these traces help evaluate the progress and quality of the run.
There’s another chart besides the trace charts, and it starts off full of useful data. This is the scatter plot in the lower right, labelled “Mutation Count vs. Date”. It compares the date of each uploaded sequence with the number of mutations in that sequence, along with a linear regression. A quick glance at this will help determine whether you have outliers that you may want to remove from your data. For example, in the example below, it’s clear one of these dots is not like the others:
Once the run actually starts, things get much more interesting. Note that it’s way more fun to do it yourself: just go to delphy.bio, select the “Ebola: Gire et al 2014” demo, click the “Run this demo” button, and then click the run button, like this:
Behind the scenes, Delphy starts generating base trees. Every time a new base tree is sampled, the phylogenetic tree changes shape and each trace chart updates with a new datapoint. Beneath each of seismograph-style charts, the distribution updates to include the new value. In an older version of Delphy, we showed each new tree as it was generated. It helps express how the process works, but it’s not actually that useful, so we took it out to make space for better trace charts.
There’s a general pattern to how the phylogenetic tree on the left updates during a run: it changes dramatically for a bit before it settles down. Recall that the goal of the algorithm is to create a bunch of equally plausible sample trees. The first tree you get is pretty good, but it’s still somewhat random. It can take a bit of the run to reach an equilibrium. This is reflected in the how much the phylogenetic tree changes with each sample: as the run approaches an equilibrium, the phylogenetic tree becomes more stable 2.
As we get more and more of the sample trees, we want to see what parts of the tree show up consistently. We use a “Maximum Clade Credibility” tree 3,4, one of the conventional ways to show those common features for data like this. Basically, every branching point is evaluated to see how many of the base trees it appears in. The ones that appear in 90% or more5 of the base trees are colored darker.
With Delphy, our goal was to create transparency into all the data behind the phylogenetic tree. During the run phase of Delphy, having visibility into all the traces can help experts get fine grained detail on the run’s progress. For non-experts, we also have a simpler way to tell how the run is progressing. In the top right of the tab, there’s a progress meter indicating how reliable or robust the run is so far. This summarizes the trace data and tells you whether it’s too soon to start exploring, good enough to vet your data, robust for exploration, or so stable that it’s really unlikely to change any more.
Our recommendation is to pause the run once the results are stable. Then open up the other tabs to take an initial look at the results. This can highlight bad data in the sequences that were uploaded to Delphy. If you find some, it could be worthwhile to remove that sequence and start another run–and because Delphy is so quick, that’s not a costly thing to do. But if everything looks good, then you can go back to the trees tab, unpause the run, and wait a few more minutes for the results to become robust. We’re hoping that having such a quick and easy tool will encourage more people to explore phylogenetic data on their own. In the next part of the series, we’ll walk through the sorts of analysis you can do in Delphy.
We’d love to hear what you’re working on, what you’re curious about, and what messy data problems we can help you solve. Drop us a line at hello@fathom.info, or you can subscribe to our newsletter for updates.