Background Science
Cellulose, enzymes, and directed evolution
The biology behind this project — from cellulose's molecular structure to how we use laboratory evolution to engineer new protein functions.
The substrate
What is cellulose, and why is it hard to break down?
Cellulose is the structural polymer that makes up plant cell walls. Chemically, it is a linear chain of glucose molecules linked by β-1,4 glycosidic bonds — the same building blocks as starch, but connected at a different angle. That seemingly small chemical difference is what makes cellulose so recalcitrant.
Where starch forms loose, helical coils that are easily digested, cellulose chains pack tightly into hydrogen-bonded, crystalline microfibrils. This crystalline structure is extraordinarily stable — it resists most chemical and enzymatic attack, which is why wood is strong and cotton doesn't dissolve in water.
To fully convert cellulose to glucose requires breaking through multiple levels of structural complexity: first loosening the crystalline packing, then fragmenting the chains, then cleaving the resulting disaccharide cellobiose into free glucose. Each step requires a different enzyme class, acting synergistically.
The enzyme system
Two real activities, and a third we're trying to add
No single natural enzyme converts crystalline cellulose to glucose alone — it normally takes three activities working in relay. CxnA already provides two of them, fused into one protein. The third is what this project is testing.
Exoglucanase
Confirmed in CxnA
A dual-specificity exoglucanase/xylanase: attacks cellulose and xylan from chain ends, releasing the disaccharide cellobiose.
Endoglucanase
Confirmed in CxnA
Cleaves β-1,4 bonds at internal positions within the cellulose chain, not at the ends. This creates new chain ends for Cex to attack, making the two activities strongly synergistic.
Our starting enzyme
CxnA: an engineered bifunctional cellulase
CxnA isn't a protein nature produces on its own. It was engineered by Duedu & French at the University of Edinburgh in 2016 [1], by fusing the catalytic domains of two separate cellulose-degrading enzymes naturally produced by Cellulomonas fimi — a Gram-positive soil bacterium known for an unusually effective natural cellulose-degradation system — into a single polypeptide chain.
What makes CxnA notable is that the fusion genuinely works: the resulting protein shows both exoglucanase and endoglucanase activity in one chain, something most organisms only achieve by deploying two separate enzymes. This project asks whether the same strategy can go one step further — not by adding a third domain, but by evolving a third activity directly into one of the two domains CxnA already has.
CxnA has exactly two domains: Cex (family GH10, exoglucanase/xylanase) and CenA (family GH6, endoglucanase), joined through a carbohydrate-binding module (CBM). There is no third, pre-existing domain. The directed-evolution target in this project is to evolve new β-glucosidase-like activity directly within Cex's own active site. Cex (GH10) and the β-glucosidases of family GH1 both belong to clan GH-A, sharing a (β/α)₈ fold and an equivalent pair of catalytic glutamates — which is what makes coaxing glucosidase chemistry out of Cex's own active site structurally plausible, if ambitious: specificity engineering of Cex itself has precedent [3], though not previously toward glucosidase activity.
Key datum: Structural validation
The AlphaFold3 model of the GH10/Cex domain was validated against the published crystal structure 1FHD (Notenboom et al., 1998 [3]). RMSD = 0.293 Å over aligned Cα atoms — indicating high structural accuracy and supporting the reliability of downstream docking predictions. See Results for full validation analysis.
CxnA at a glance
| Source organism | Cellulomonas fimi (origin of the two fused domains) |
|---|---|
| Built by | Duedu & French, 2016 [1] — an engineered fusion, not a natural isolate |
| Institution | University of Edinburgh |
| Domain 1 | Cex — family GH10, exoglucanase/xylanase |
| Domain 2 | CenA — family GH6, endoglucanase |
| Joined via | Carbohydrate-binding module (CBM) |
| Crystal structure ref | 1FHD (Cex, Notenboom 1998 [3]) |
| Predicted-structure RMSD vs 1FHD | 0.293 Å (validated) |
The method
What is directed evolution?
A laboratory technique that mimics natural selection — but compressed into weeks, and directed toward a specific functional goal.
Introduce random mutations
Error-prone PCR (epPCR) amplifies the target gene with a deliberately elevated error rate, creating a library of sequence variants — each a slightly different version of the protein. In principle this can generate libraries of thousands of variants; what this project actually recovered was far smaller — see Results for the real number and why. This is analogous to natural mutation, but we control the rate.
Screen colonies on indicator plates
Every colony in the library — not just survivors of a growth challenge, an approach we tested first and found too weak to enrich effectively — is spread onto plates carrying a colour-change substrate: Congo red for retained endoglucanase activity, X-gluc for the β-glucosidase-like activity being evolved. This narrows the library down to a shortlist worth quantifying.
Quantify the shortlist
Shortlisted colonies are grown in liquid culture and lysed, then assayed against a fluorogenic substrate (MUC) to put an actual number on activity rather than relying on a plate colour by eye. This is the step that turns "looked darker on the plate" into a comparable signal.
Sequence and verify
Hits are sequenced to identify which mutations they actually carry. Because epPCR mutagenesis is random, not targeted, sequencing is essential — a colony can look promising on a plate for reasons that have nothing to do with the mutation dry-lab modelling would have predicted.
Why directed evolution works
Natural evolution explores protein sequence space over millions of generations and years. Directed evolution recapitulates this process in the lab by: deliberately narrowing the sequence space explored (we start from a protein that already has related structure), applying a reporter-based screen to find rare interesting variants directly rather than waiting for natural selection, and iterating rapidly (weeks, not millennia).
The key insight is that you don't need to rationally design every mutation. The screening step acts as the filter — you just need to generate enough sequence diversity that beneficial mutations exist somewhere in the library, and then find them.
This project also uses computational structural analysis (AlphaFold3 + docking + FoldX) to identify which residues are worth paying attention to — but epPCR itself mutates the whole gene at random. The computational work can't steer where a mutation lands; it can only flag candidate positions (like Cex N44W or CenA K671S) worth prioritising when deciding what to sequence and follow up on afterward.
Important limitation
Computational predictions guide analysis but cannot replace experimental testing, and in this project they did not dictate which mutations the wet-lab library actually carried — epPCR is random, not targeted. Docking scores predict approximate binding geometry; they do not predict whether a mutation will actually confer catalytic activity in a living cell. See Results for how the dry-lab predictions and the actual wet-lab mutations compared.
Scientific foundation
Key literature this project builds on
Duedu & French — building CxnA
The original construction of CxnA: fusing the Cex (GH10) and CenA (GH6) catalytic domains from Cellulomonas fimi into one engineered bifunctional polypeptide, and providing the sequence and expression system this project uses as its starting point.
Notenboom et al. — GH10/Cex crystal structure
Crystal structure determination of the Cex exoglucanase domain (PDB: 1FHD and related entries), establishing the active-site architecture including the catalytic glutamates. [3]
Damude et al.; Cockburn et al. — CenA/GH6 mechanism
Characterisation of GH6-family endoglucanase catalysis, including the catalytic residue architecture used to annotate CenA's active site. [4][5]
In-text citations on this page jump to the summaries above. The complete bibliography, including references not discussed in depth here, is on the Results page.