micro NovaSeq Run #4
Pool was mailed to the University of Colorado Genomics and MicroArray Core at University of Colorado on 01-27-2021 promising sequencing to be run on Monday, Feb. 1st Monday, Feb. 8th. Feb. 22nd
QC from University of Colorado Genomics Core
Sequencing Report From University of Colorado Genomics Core:
Flowcell Summary
Clusters (Raw) | Clusters(PF) | Yield (MBases) |
|---|
Clusters (Raw) | Clusters(PF) | Yield (MBases) |
|---|---|---|
1,276,674,048 | 956,877,735 | 499,490 |
Lane Summary
Lane | PF Clusters | % of the | % Perfect | % One mismatch | Yield (Mbases) | % PF | % >= Q30 | Mean Quality |
|---|
Lane | PF Clusters | % of the | % Perfect | % One mismatch | Yield (Mbases) | % PF | % >= Q30 | Mean Quality |
|---|---|---|---|---|---|---|---|---|
1 | 486,226,470 | 100.00 | 100.00 | NaN | 253,810 | 76.17 | 89.20 | 35.01 |
2 | 470,651,265 | 100.00 | 100.00 | NaN | 245,680 | 73.73 | 87.99 | 34.77 |
This will use the 1-step PCR primers. All of the adaptor sequence is added at once. This saves on time and reagents.
Include Yoon plate for resequencing
Comments
I would like them as soon as possible ideally. Before break is the cutoff if I am going to get it out early January as promised to Paul. What is your quickest timeline? There is plenty of room still in the run.
Hi Gregg, the ~160 samples are already extracted and I will arrange with Martina to submit them asap, and the 24 samples should be extracted within the next week or so, so by the end of next week likely.
Forgot to confirm; the $25 per sample is for both the library prep and sequencing? THanks!
For now, yes it does as long as the samples are being bundled on an EPSCoR NovaSeq run. iSeq runs are priced per run and any completely separate NovaSeq runs will have to be dealt with if they ever come up. Prices will be changing when we implement the new mandated pricing structure.
Checking in regarding sample barcoding/ID of the ones that Martina submitted before break. She is no longer working for us (we have not been able to reach her since the break), so I’m trying to get the excel datasheet (sample onboarding “EPSCoRLibrary4SampleInput” ) updated myself. Could we come pick up the boxes of the DNA extracts so we can fill out the datasheet with the sampleIDs corresponding to the barcodes? On another note, with the delays in sequencing turnaround can we still add samples to this run4 to be prepped by you?
We will fill in the samples in the order they were added to the plates and ping you when the information we do not know will need to be added. How many samples are we talking about adding? We were hoping to send it out early next week to whichever facility can promise the shortest turn around with the Illumina reagent backlog. Do you want us to return the tubes to you or store them here?
No worries about the additional samples, it would be a while before we could submit them, so you ahead and send this plate off when ready so we can get results back. Yes, I would like to get the tubes back as we will be using them for qPCR analysis. Would I be able to pick them up Thursday morning by any chance?
Hi @Gregg Randolph & @Shannon Harris, I’m unsure of the best place to ask this question, so I’ll just throw it up on this library prep that contained some of my resequenced samples. It is my understanding that the concentration of each PCR product is checked before creating the equimolar library. Where could I expect find the spot-checked DNA concentrations for the following samples? Thanks.
G152SG G011SG G092SG G165SG G049SG G138SG G023SG G127SG G107SG G082SG G095SG G166SG G065SG G174SG
I’ve found this folder “//project/microbiome/data/seq/Concentrations” but it only has concentrations for novaseq1 and novaseq2.
For NovaSeq4 and one step moving forward, we only spot check with qPCR our plates (via one column for NS4). We only adjust the pooling ratios if there is a consistent order of magnitude difference between plates. Otherwise, 2 ul from each sample is added to each pool. I can check if those samples were qPCRed. If not, we can do that if it is needed.
Hey Greg, thanks for the info. For a little background, I’m trying to better understand why there is such discrepancy in sequencing depth among my samples. Am I understanding this correctly that PCRs of individual samples are not verified and that the concentration of all samples on a plate is checked by spot checking samples from a single column on that plate?
I can get you your starting concentrations for those samples relatively easily. All sample concentrations are checked and normalized to 10 ng/ul prior to PCR. If those samples were less than 10 ng/ul at the start, then they entered PCR at their starting concentrations.
Ok, thanks. That would be helpful. Out of curiosity, how are we ensuring a sample amplified? I’m concerned as there are many of my samples that were < 3000 reads, with a good number under 500, especially when other samples frrom the same project had upwards of 200k.
Hi Gordon. Are the low counts for both replicates?
As for ensuring that there are enough reads, that’s mostly up to the user to determine and let us know about, as we often do not know the metadata associated with a sample. For some micro projects we have redone samples with low read counts.
Hi Alex and Greg,
Yes, some of the samples failed for both replicates and others failed for just one. In the cases where one worked, I can just use that replicate. Ok, I will have to revisit my tables and figure out which samples should be resequenced.
I have a couple of questions about the sample prep process though.
Each replicate is a single PCR, correct? Why did we chose to do a single PCR instead of triplicate which seems to be the norm for amplicon preps? Likewise, why did we not standardize samples and create an equimolar library prior to sequencing? In my previous experience, this step seems to be a required step and may account for the discrepancy in read depth among samples.
Hi Gordon. For each sample, we replicate the PCR prep in an entirely separate prep, with unique MIDs, so that users can see the extent to which they concur. Rather than triplicate reactions, we do duplicates and have had good success, with high correlation in the read counts per taxon between duplicates. In contrast, I believe it is common practice to use the same MID in triplicate and to pool the products for sequencing, which gives no information about the value of replicates.
As for quantitating the DNA and making it equimolar, we do this (see @Gregg Randolph's comment above). But if samples are below 10 ng/µL, we do not dilute the sample.
Hi Alex, thanks for the info. The duplicate, separate PCRs make sense, and the info gained from keeping them separate is certainly useful. You are correct, what I’ve seen done in the past is to combine three separate PCRs using a single mid combo in order to minimize individual PCR bias. As long as our replicates are highly correlated, I assume this is a non-issue. However, has anyone encountered a situation in which the replicates were not correlated? I’ve had a few like this, but it commonly resulted from a failed PCR/failed sequencing in one of the samples.
As for the standardization to equimolar concentrations, I interpreted Gregg’s comment above to mean the DNA template was standardized to 10ng/uL before the PCR was performed, not the PCR product. If this is the case, couldn’t we expect significant variation in the concentrations of PCR product based upon reaction efficiency, template specificity, and inhibitor concentration etc.?
Hi Gordon. Sorry, I thought you when you wrote ‘samples’ you meant ‘templates’, not ‘products’, but now I see that later you refer to equimolar libraries. It is absolutely true that different PCRs could give you different amounts of product. I believe we spot check each plate (I believe it is one column of 8 samples) and pool the plates accordingly.
If we have user feedback that this is insufficient for their samples, we could do more quantitation. However, I suspect that a lot of previous labs did this because they did not use a model for read counts (multinomial). The multinomial can readily accommodate different levels of sampling (read counts).
Yes, Alex is correct above. We spot checked 8 samples per plate. We only adjust the pooling ratio if the plates are different by a consistent order of magnitude or more. We do not adjust if there is a high or low outlier in the tested column.
Hey Alex, no worries. Thanks for the help in deciphering everything thats going on with the sample prep. I see your point about the use of the multinomial to account for discrepancy in per sample read depth. However, I’m cautious to accept this at this time. Part of me says, that samples with widely differing sequencing depths (some that span 4/5 orders of magnitude) are not generated by the same processes. Are you aware of information regarding the effect of vastly different sequencing depths/counts on the model’s ability to estimate proportions, especially for the poorly sequenced samples? It is my understanding that samples with higher read counts are given more weight. At what point do my samples with low read counts become overwhelmed by the observed counts of the samples with vastly higher read depths? In an ideal world, all of my samples would have 50k reads. I know this is highly unlikely. However, I want to understand at what point the differences in sequencing depths matter.
Gordon, you are mixing a few different points, but I hear your concern.
If you think that the two replicate PCR preps from the same template are not the result of the same process (PCR), it seems unlikely that statistics or concentration adjustment of templates is going to help escape that problem. What do you have in mind?
As for combining information from replicates or multiple samples in a model, and worrying about things overwhelming others, I would need greater clarity about the model you have in mind. That said, in CNVRG we propagate our uncertainty among parameters, and in general samples with little information (few reads) contribute less to parameter estimates. That seems like a good idea.
As for the returns on greater sequencing depth, for a single sample, this would be straightforward to illustrate. Here is a figure for the uncertainty on the frequency of a taxon that has frequency 0.5, or one that has frequency of 0.05 (both pretty large). A frequency of 0.5 has the maximum uncertainty of any possible frequency, so is an upper bound.
Thanks for the reply. I feel that we may not be understanding each other entirely. I will try to put something together and circle back.
Hi Gregg, is there still time and space to add samples to this run #4? If our project gets approved today, I would submit ~160 DNA extracts, and would have an additional 24 samples from another student project outside of the Micro Project but from UW; could those be included as well, and if so, what would the costs be to include those ‘outside’ samples? THanks!