Results
What we found — and what we didn't
CxnA is a two-domain enzyme: Cex (an exoglucanase) and CenA (an endoglucanase), fused together. We used error-prone PCR to scramble its DNA at random, then screened the resulting mutants to see whether any of them had picked up a useful new or improved activity. Computational predictions guided where we looked; the mutagenesis itself was random. No mutant clearly beat the unmodified enzyme — but the screening pipeline we built to find out is itself the result, and it's reusable for the next round.
Results status summary
Dry Lab
Wet Lab
Computational results
Dry Lab — where should directed evolution even look?
Before touching a single bacterium, we used structure prediction and docking software to make an educated guess at which parts of CxnA were worth targeting — and, just as importantly, which predictions we should and shouldn't trust.
What does CxnA actually look like?
CxnA has no solved crystal structure of its own, so we predicted one with ColabFold/AlphaFold2 — colouring the result by domain: Cex (exoglucanase, green), the carbohydrate-binding module that holds the two halves together (blue), and CenA (endoglucanase, pink), joined by flexible proline/threonine-rich linkers.
Checking Cex's active site against a real crystal structure
Unlike the fusion as a whole, Cex alone has a published crystal structure to check against: 1FHD, a Cex-family enzyme solved with a bound inhibitor [2]. We superimposed our predicted Cex structure onto it to see how well the model's active site matches reality.
align (iterative outlier rejection) over 1,958 matched atoms
against the 1FHD crystal structure — the converged, best-fitting core rather than a
pure Cα figure (1,958 exceeds this domain's actual Cα count). A separate, full-domain
Cα-only superposition, run live in the interactive viewer below, gives 0.579 Å over
311 pairs; the two numbers measure different things and aren't directly comparable to
each other, or to the crystallographers' own crystal-form comparison, which used yet
another alignment definition. Either way, the active-site geometry, including the
catalytic residues GLU127 and GLU233, is reproduced faithfully.
Interactive 3D overlay: predicted Cex structure (teal) superimposed on the 1FHD crystal structure (orange) — Cα RMSD 0.579 Å over 311 matched pairs in this particular alignment run. Yellow spheres mark GLN87, GLU127 and GLU233. Drag to rotate; use the toggle buttons to view each structure independently.
Why this step matters
Every docking result later in this page depends on the predicted structure being trustworthy. This validation is what lets us treat the Cex model as a stand-in for the real enzyme when testing how it interacts with sugars computationally.
Does Cex already have a little latent β-glucosidase activity — and where would you nudge it?
Cex's fold resembles known β-glucosidase enzymes closely enough that we asked a direct question: can cellobiose — the two-sugar fragment a true β-glucosidase would split into glucose — fit into Cex's existing active site at all? We docked cellobiose into the validated Cex model with AutoDock Vina and analysed the predicted contacts with PLIP.
Multiple-sequence alignment against 15 homologous GH10 enzymes, cross-checked with ConSurf conservation scoring, confirmed that the catalytic motifs around GLU127 (an unusual "NEA" variant of the normal "NEP" motif) and GLU233 (the "ITELD" motif) are essentially universal across the family — strong evidence that both residues are functionally essential, not incidental.
Cex coloured by ConSurf conservation (red = most conserved). The NEA and TELD boxes mark the two catalytic motifs — both are deep red, i.e. essentially invariant across 15 homologous enzymes.
Where would you aim mutagenesis, if you were targeting it?
HotSpot Wizard flagged candidate residues near the active site that are mutable without obviously destroying the protein's fold. Those candidates were then screened computationally with FoldX (for stability) and re-docked with cellobiose (for binding). The most stabilising candidate to emerge was Asn44 → Trp (ΔΔG −3.2 kcal/mol, i.e. predicted more stable than wild type). Stability is not activity, so this ranks candidates by tolerability rather than by predicted gain in function.
CenA, the other half: where could the existing activity be improved?
CenA (the endoglucanase half of CxnA, family GH6) already works — the Project B goal isn't to invent a new activity but to make an existing one better. We docked cellotetraose (a four-sugar cellulose fragment) into CenA's active site and confirmed contacts with four catalytic aspartates: Asp595 (a stabiliser that fine-tunes the pKa environment around the catalytic acid, not itself required for catalysis), Asp631 (acid catalyst), Asp666 (a secondary, pKa-adjusting partner to Asp631), and Asp771 (base catalyst).
CenA coloured by ConSurf conservation across 41 homologous GH6 sequences (magenta = most conserved, teal = most variable). 13 of the 27 residues contacting the docked substrate are fully conserved — a strong signal of functional importance.
As with Cex, candidate mutations were screened by combining FoldX stability predictions with re-docking scores — this time checking both metrics together, plus whether the catalytic geometry held up across repeated modelling runs.
An important honesty note: prediction guided priorities, it didn't dictate the mutants
Read this before the wet-lab results
N44W and K671S are the dry lab's top computational candidates — but the actual wet-lab mutant library wasn't built by inserting those specific changes. It was built with error-prone PCR, which introduces mutations randomly across the gene. The computational work told us the Cex and CenA active sites were worth targeting and gave us a reference point for interpreting what random mutagenesis turned up — it didn't guarantee any particular mutation would actually appear in the library. Keep that distinction in mind in the section below: none of the sequenced wet-lab mutants carry N44W or K671S.
Experimental results
Wet Lab — building and screening real mutants
Three screens, run in sequence, each narrowing down which of the random mutants were worth a closer look: a visual test for endoglucanase activity, a colour test for glucosidase activity, and finally a quantitative fluorescence assay we had to build and validate from scratch.
Before the library: a growth-based selection, tried and set aside
Before building and screening a random mutant library, we first asked whether growth on cellobiose as the sole carbon source could select for improved β-glucosidase-like activity directly — pool a library, grow it on cellobiose over repeated rounds, and let only the active variants multiply, without screening any colonies by hand. E. coli expressing wild-type CxnA was compared against the same strain expressing a secreted β-glucosidase as a positive control, each grown in parallel on glucose (a fitness control) and on cellobiose (the actual test).
(A) OD₆₀₀ at 24 h. The two strains are indistinguishable on glucose (1.82 vs 1.87), ruling out a general fitness difference. On cellobiose, the β-glucosidase control reaches 0.53 against 0.23 for CxnA — the latter close to the casamino-acid-only background, confirming CxnA itself doesn't support growth on cellobiose. (B) Viable counts by serial dilution plating tell the same story independently. (C) Cumulative enrichment over repeated rounds of liquid selection, against the roughly 10,000-fold enrichment a rare variant would need to become recoverable.
The plasmid, and an uninvited stowaway
CxnA was carried on a pSB1C3 plasmid backbone with a lac promoter and ribosome binding site ahead of the insert. Transforming E. coli with this construct retained both exo- and endo-glucanase activity, confirming the starting enzyme was functional before any mutagenesis began.
Plasmid map of the CxnA construct. CxnA itself is only two domains — Cex (exoglucanase) and CenA (endoglucanase) — joined through a carbohydrate-binding module (CBM). The duplicated chloramphenicol-resistance region on the right was an unplanned discovery (below).
Building the mutant library — smaller than planned
Error-prone PCR (epPCR) was applied only to the CxnA insert, using manganese chloride to deliberately lower the copying polymerase's accuracy — the backbone, promoter and ribosome binding site were amplified separately with a high-fidelity polymerase so mutations wouldn't land somewhere that would confound the activity readout. We ran this on two versions of the gene in parallel: the native sequence, and a codon-optimised version ordered as synthetic DNA blocks (to work around the gene's unusually high, amplification-hostile GC content — 71%).
24 codon-optimised and 15 native-sequence colonies were recovered. Fifteen of the 24 codon-optimised colonies were picked forward, giving the 30 mutants screened below.
Screen 1 — Congo red: did endoglucanase activity survive?
Congo red dye binds intact cellulose (here, the related polymer CMC) but washes off wherever an active endoglucanase has chewed it up — so a colony with working CenA activity shows up as a clear halo on a red background. This was run first, as a cheap way to immediately discard any mutant that had simply broken the protein.
Congo red screen. (A) All 15 codon-optimised mutants, numbered. (B) All 15 native-sequence mutants (only 15 total colonies were available in that library). A clear halo = endoglucanase activity retained.
Screen 2 — X-gluc: any sign of new glucosidase activity?
X-gluc is a colourless compound that turns a colony visibly blue if the cell is producing β-glucosidase-like activity. Every mutant from both libraries — 30 colonies total — was screened this way in parallel with Congo red, alongside controls: a known β-glucosidase (positive control), wild-type CxnA, and a non-functional His-tagged CxnA (negative control).
Controls for both screens. Left: Congo red. Right: X-gluc — note that wild-type CxnA (top right spot) already shows a faint blue tint, not just the positive-control β-glucosidase (bottom left).
X-gluc screen on all 30 mutants. (A) codon-optimised. (B) native sequence.
Screen 3 — putting a number on it: the MUC fluorescence assay
Congo red and X-gluc are useful for sorting mutants into rough buckets, but neither gives a number you can actually compare between colonies. To measure exoglucanase activity precisely, we used MUC — a substrate that releases a fluorescent molecule (4-MU) in proportion to how much active enzyme is present. Getting trustworthy numbers out of this assay took six rounds of calibration (the full story — including a substrate mix-up that briefly looked like a real signal — is in the lab notebook); the validated version is summarised below.
Establishing a trustworthy assay. (A) Fluorescence is only proportional to 4-MU concentration below 6.25 µM — above that, the reading undershoots for optical reasons, not biology, so all rates were fit only within the reliable range. (B) Adding more lysate gave progressively less signal per microlitre — a volume artefact, not reduced activity. (C) Fixed by titrating active lysate against inactive (frameshift) lysate instead, which restored a straight-line relationship at two different temperatures/settings. (D) No signal from substrate, medium, or host alone — the frameshift control sits at just 0.7% of induced CxnA.
With the assay validated, we quantified exoglucanase activity for the seven mutants flagged by the earlier screens, each compared against a wild-type control grown and lysed alongside it on the same day. Seven variants were carried forward to the quantitative assay, limited by the capacity of a single 96-well plate once controls, blanks and standards were accommodated. The selection deliberately spanned both extremes of the Congo red screen — variants with the largest halos and variants with no clearing at all — rather than sampling only the apparent improvers. Native colony 12, which also showed no clearing, was not carried forward for this reason; T14 and T15 represent that class.
Activity of all seven flagged mutants, relative to a same-day wild-type control (bars = median of 3 technical replicates; whiskers = range). The shaded band marks how much two colonies of the same, unmutated genotype normally differ from each other just from day-to-day prep variation — only a bar that clears this band by itself is a plausible real effect.
We also tried to run this as a MUG assay
The same lysates were also assayed with what was intended to be MUG, the β-glucosidase substrate this project was trying to evolve activity toward. Every sample lit up immediately, including the dead-enzyme frameshift control. A controlled follow-up — substrate and lysate dilution series, boiled lysate, cultures grown without IPTG, and cell-free medium in place of lysate — diagnosed the cause: the substrate in hand was MUGal, a β-galactosidase substrate, not MUG. TOP10 carries the lacZΔM15 deletion and the plasmid supplies the lacZα fragment that complements it, reconstituting host β-galactosidase, which hydrolyses MUGal regardless of what CxnA is doing. The assay was discontinued once this was confirmed. It does not affect the MUC results above: the frameshift control sits below 1% of induced signal there, so this background enzyme does not touch the cellobioside substrate. Full troubleshooting: 15 September notebook entry.
What mutations were actually there?
Eight samples were sent for Sanger sequencing (MRC PPU DNA Sequencing and Services): the seven variants carried through the quantitative screen, plus the unmutated codon-optimised construct as a parental reference. Plasmid DNA was isolated by miniprep (Qiagen) and quantified on a NanoDrop 2000 (Table 2).
All four codon-optimised variants — CO2, CO4, CO7 and CO12 — returned the parent sequence with no mutations detected. Sequencing failed for T14 and T15, so neither of the two dead variants can be characterised genetically. Of the seven variants, only T4 returned usable evidence of a mutation: a single point substitution in each domain, changing residue 14 from GGC to GAC (glycine → aspartic acid) in the exoglucanase domain, and residue 614 from aspartic acid to glycine in the endoglucanase domain. Changes flagged elsewhere in the chromatograms were consistent with signal decay rather than real substitutions.
Read-length limits left an unsequenced gap in the middle of every read. The gap spans the carbohydrate-binding module in all samples and extends into the Cex or CenA coding sequence in some, and its extent differs from sample to sample — so coverage is not uniform, and a mutation falling inside the gap would not have been detected.
It also shows the 1.40× threshold rests on too small a sample. That figure came from a 9.5% CV measured between two colonies each of two genotypes. These four independently picked isolates, grown and assayed in the same preparation, span 0.57× to 1.39× — roughly four times wider — and the lowest of them falls outside the ±3σ interval the 9.5% figure implies. Both estimates rest on a handful of observations, but the wider one is measured across the comparison the screen actually makes: different picked colonies, not replicate colonies of one clone. On that basis a real improvement would need to be considerably larger than 1.4× before a single-colony measurement could resolve it.
Coverage is incomplete, so "no mutations detected" is not the same as "no mutations" — see above for where the unsequenced gap falls. Taken as a bound rather than a proof: over the regions that were read these four isolates are identical to the parent, and the simplest reading of a 0.57–1.39× spread among them is preparation and colony variation rather than four separate mutations that all happened to land in the unread gap.
So — did it work?
No mutant demonstrated a confirmed improvement in activity. CO12, the only variant to read above wild type, returned the parent sequence — placing its 1.35× average inside the range spanned by genetically identical isolates, so it is not a candidate for a further round. The screen's useful output is a quantified noise floor rather than a hit: variation between preparations is wide enough that a real improvement would have to be considerably larger than 1.4× before a single-colony measurement could resolve it. Two variants, T14 and T15, lost function entirely; sequencing failed for both, so the explanation rests on phenotype — simultaneous loss of two activities whose active sites lie several hundred residues apart is more easily explained by loss of the full-length protein than by two independent active-site mutations.
What did work: we built a complete, repeatable pipeline — random mutagenesis, two fast qualitative screens, a validated quantitative fluorescence assay, and sequencing to close the loop — that took a library from 30 random colonies to a quantified detection limit and one genetically characterised loss-of-function variant, in a single round. The library was small and the GC-rich gene was a genuine technical obstacle throughout; a larger library run through the same pipeline is the obvious next step, not a different approach.
Two observations point the same way about the codon-optimised library. All four of its variants that were sequenced returned the parent sequence, and all 15 of its colonies retained a Congo red halo — against 7 of 15 for the native-sequence library. The simplest reading is that error-prone PCR introduced few or no mutations into the codon-optimised construct, consistent with the gBlock-and-Gibson route reconstituting parental sequence, and that the native library is where this project's real variants are: T4, with a substitution in each domain, and T14 and T15, both of which lost function entirely. Only four of the 15 codon-optimised colonies were sequenced, so this is the most parsimonious explanation rather than an established one.
Why we're reporting a null result at all
A "nothing improved, but here's exactly why and how confident we are" result is a real scientific outcome, not a failure to hide. Every negative or inconclusive finding on this page — the T14/T15 loss-of-function variants, CO12 turning out to be a parental isolate rather than a hit, the small library sizes — is reported as measured, not rounded up or omitted.
Bibliography
Full reference list
All sources cited across the project website, matching the final project report.
- Notenboom, V., Birsan, C., Warren, R.A.J., Withers, S.G. & Rose, D.R. (1998). Exploring the cellulose/xylan specificity of the β-1,4-glycanase Cex from Cellulomonas fimi through crystallography and mutation. Biochemistry, 37(14), 4751–4758. doi:10.1021/bi9729211
- Notenboom, V., Williams, S.J., Hoos, R., Withers, S.G. & Rose, D.R. (2000). Detailed structural analysis of glycosidase/inhibitor interactions: complexes of Cex from Cellulomonas fimi with xylobiose-derived aza-sugars. Biochemistry, 39(37), 11553–11563. PDB: 1FHD. doi:10.1021/bi0010625
- Duedu, K.O. & French, C.E. (2016). Characterization of a Cellulomonas fimi exoglucanase/xylanase–endoglucanase gene fusion which improves microbial degradation of cellulosic biomass. Enzyme and Microbial Technology, 93–94, 113–121. doi:10.1016/j.enzmictec.2016.08.005
- Damude, H.G., Ferro, V., Withers, S.G. & Warren, R. (1996). Substrate specificity of endoglucanase A from Cellulomonas fimi: fundamental differences between endoglucanases and exoglucanases from family 6. Biochemical Journal, 315(2), 467–472. doi:10.1042/bj3150467
- Cockburn, D.W., Vandenende, C. & Clarke, A.J. (2010). Modulating the pH–activity profile of cellulase by substitution: replacing the general base catalyst aspartate with cysteinesulfinate in cellulase A from Cellulomonas fimi. Biochemistry, 49(9), 2042–2050. doi:10.1021/bi1000596
- Adasme, M.F. et al. (2021). PLIP 2021: expanding the scope of the protein–ligand interaction profiler to DNA and RNA. Nucleic Acids Research, 49(W1), W530–W534. doi:10.1093/nar/gkab294
- Ashkenazy, H. et al. (2016). ConSurf 2016: an improved methodology to estimate and visualize evolutionary conservation in macromolecules. Nucleic Acids Research, 44(W1), W344–W350. doi:10.1093/nar/gkw408
- Eberhardt, J., Santos-Martins, D., Tillack, A.F. & Forli, S. (2021). AutoDock Vina 1.2.0: new docking methods, expanded force field, and Python bindings. Journal of Chemical Information and Modeling, 61(8), 3891–3898.
- Schymkowitz, J. et al. (2005). The FoldX web server: an online force field. Nucleic Acids Research, 33(suppl_2), W382–W388. doi:10.1093/nar/gki387
- Beckman, R.A., Mildvan, A.S. & Loeb, L.A. (1985). On the fidelity of DNA replication: manganese mutagenesis in vitro. Biochemistry, 24(21), 5810–5817. doi:10.1021/bi00342a019
- Olofsson, K., Bertilsson, M. & Lidén, G. (2008). A short review on SSF — an interesting process option for ethanol production from lignocellulosic feedstocks. Biotechnology for Biofuels, 1(1). doi:10.1186/1754-6834-1-7
- Zajki-Zechmeister, K. et al. (2021). Processive enzymes kept on a leash: how cellulase activity in multienzyme complexes directs nanoscale deconstruction of cellulose. ACS Catalysis, 11(21), 13530–13542. doi:10.1021/acscatal.1c03465
- Ejaz, U., Sohail, M. & Ghanemi, A. (2021). Cellulases: from bioactivity to a variety of industrial applications. Biomimetics, 6(3), 44. doi:10.3390/biomimetics6030044
- Zadorozhny, A.V. et al. (2025). Cellulases: key properties, natural sources, and industrial applications. Vavilov Journal of Genetics and Breeding, 29(8), 1348–1360. doi:10.18699/vjgb-25-141