Background Science

Cellulose, enzymes, and directed evolution

The biology behind this project — from cellulose's molecular structure to how we use laboratory evolution to engineer new protein functions.

The substrate

What is cellulose, and why is it hard to break down?

Cellulose is the structural polymer that makes up plant cell walls. Chemically, it is a linear chain of glucose molecules linked by β-1,4 glycosidic bonds — the same building blocks as starch, but connected at a different angle. That seemingly small chemical difference is what makes cellulose so recalcitrant.

Where starch forms loose, helical coils that are easily digested, cellulose chains pack tightly into hydrogen-bonded, crystalline microfibrils. This crystalline structure is extraordinarily stable — it resists most chemical and enzymatic attack, which is why wood is strong and cotton doesn't dissolve in water.

To fully convert cellulose to glucose requires breaking through multiple levels of structural complexity: first loosening the crystalline packing, then fragmenting the chains, then cleaving the resulting disaccharide cellobiose into free glucose. Each step requires a different enzyme class, acting synergistically.

Why does the β-1,4 linkage matter so much?

The β configuration at the glycosidic bond causes each glucose unit to rotate 180° relative to its neighbour, producing an extended, flat ribbon rather than a helix. These flat ribbons can stack and hydrogen-bond to each other very efficiently, forming the tight crystalline bundles that give cellulose its mechanical strength.

Most animals cannot make cellulases, which is why cows and many other herbivores depend on gut microbes to break β-1,4 bonds for them; a few, such as termites, produce endogenous cellulases of their own alongside their symbionts. Fungi and bacteria such as Cellulomonas fimi (the source organism of CxnA) are among the most efficient natural cellulose degraders known.

Cellulose lattice: hexagons represent glucose units, linked in-chain by β-1,4 bonds and between chains by hydrogen bonds; one hydrogen bond is highlighted as the reason the lattice resists breaking down ⋯ ⋯ The one bond that matters One of thousands — each one weak alone, but together they lock the chains into a crystal that resists breaking down. = one glucose unit β-1,4 bond (within a chain) hydrogen bond (between chains)
Hexagons represent glucose units. Solid lines are the covalent β-1,4 bonds within each chain; dashed lines are the hydrogen bonds between chains — individually weak, collectively what makes the lattice resist breaking down.

The enzyme system

Two real activities, and a third we're trying to add

No single natural enzyme converts crystalline cellulose to glucose alone — it normally takes three activities working in relay. CxnA already provides two of them, fused into one protein. The third is what this project is testing.

Cellulose degradation relay: CenA cleaves internal bonds, Cex releases cellobiose from chain ends, and an attempted third step would split cellobiose into glucose Crystalline cellulose tightly packed, resistant CenA Shorter fragments new chain ends exposed Cex Cellobiose (two linked glucose) one step from glucose ? Glucose fermentable sugar Confirmed — CxnA already does this Attempted — see Results for what actually happened
CenA and Cex are two separate, real domains that CxnA already fuses into one protein — not a single domain doing double duty. The last step has no dedicated domain in CxnA at all; the target is to evolve that activity directly inside Cex's own active site.
Cex — family GH10

Exoglucanase

Confirmed in CxnA

A dual-specificity exoglucanase/xylanase: attacks cellulose and xylan from chain ends, releasing the disaccharide cellobiose.

Structural and mechanistic detail

Cex employs a retaining double-displacement mechanism, with two conserved glutamates — GLU233 as nucleophile and GLU127 as acid/base — coordinating catalysis. [3] It sits in an open (β/α)₈ cleft characteristic of family GH10, and is a dual-specificity enzyme: in Cellulomonas fimi it acts on xylan as well as on cellulose, and is readily assayed on soluble aryl glycosides such as 4-methylumbelliferyl cellobioside. This dual specificity is why the available crystal structures of Cex are complexes with xylobiose-derived ligands, and it is part of why the active site was considered a plausible starting point for coaxing out glucosidase-like chemistry.

CenA — family GH6

Endoglucanase

Confirmed in CxnA

Cleaves β-1,4 bonds at internal positions within the cellulose chain, not at the ends. This creates new chain ends for Cex to attack, making the two activities strongly synergistic.

Structural and mechanistic detail

Unlike Cex's retaining mechanism, family GH6 enzymes like CenA are inverting. CenA's catalytic residues — Asp595 (a stabiliser that fine-tunes the pKa environment around the catalytic acid, not itself required for catalysis on its own), Asp631 (acid catalyst), Asp666 (a secondary, pKa-adjusting partner to Asp631), and Asp771, generally assigned as the catalytic base, though the identity of the base in family GH6 has been debated — sit in an open active-site cleft rather than a tunnel, consistent with binding internal positions on a chain rather than only chain ends. [4][5] CxnA fuses Cex and CenA into a single polypeptide, which is the starting point for this project. [1]

The missing step: β-glucosidase-like activity

Converting cellobiose into two free glucose molecules is the one job CxnA's wild-type sequence doesn't do. There's no third domain to evolve here — the target is to coax this activity directly out of Cex's own active site, which structurally resembles known β-glucosidases closely enough to make the attempt plausible. Whether it worked is an experimental question, not a given — see Results for the dry-lab reasoning and the verdict.

Our starting enzyme

CxnA: an engineered bifunctional cellulase

CxnA isn't a protein nature produces on its own. It was engineered by Duedu & French at the University of Edinburgh in 2016 [1], by fusing the catalytic domains of two separate cellulose-degrading enzymes naturally produced by Cellulomonas fimi — a Gram-positive soil bacterium known for an unusually effective natural cellulose-degradation system — into a single polypeptide chain.

What makes CxnA notable is that the fusion genuinely works: the resulting protein shows both exoglucanase and endoglucanase activity in one chain, something most organisms only achieve by deploying two separate enzymes. This project asks whether the same strategy can go one step further — not by adding a third domain, but by evolving a third activity directly into one of the two domains CxnA already has.

CxnA has exactly two domains: Cex (family GH10, exoglucanase/xylanase) and CenA (family GH6, endoglucanase), joined through a carbohydrate-binding module (CBM). There is no third, pre-existing domain. The directed-evolution target in this project is to evolve new β-glucosidase-like activity directly within Cex's own active site. Cex (GH10) and the β-glucosidases of family GH1 both belong to clan GH-A, sharing a (β/α)₈ fold and an equivalent pair of catalytic glutamates — which is what makes coaxing glucosidase chemistry out of Cex's own active site structurally plausible, if ambitious: specificity engineering of Cex itself has precedent [3], though not previously toward glucosidase activity.

Key datum: Structural validation

The AlphaFold3 model of the GH10/Cex domain was validated against the published crystal structure 1FHD (Notenboom et al., 1998 [3]). RMSD = 0.293 Å over aligned Cα atoms — indicating high structural accuracy and supporting the reliability of downstream docking predictions. See Results for full validation analysis.

CxnA at a glance

Source organismCellulomonas fimi (origin of the two fused domains)
Built byDuedu & French, 2016 [1] — an engineered fusion, not a natural isolate
InstitutionUniversity of Edinburgh
Domain 1Cex — family GH10, exoglucanase/xylanase
Domain 2CenA — family GH6, endoglucanase
Joined viaCarbohydrate-binding module (CBM)
Crystal structure ref1FHD (Cex, Notenboom 1998 [3])
Predicted-structure RMSD vs 1FHD0.293 Å (validated)

The method

What is directed evolution?

A laboratory technique that mimics natural selection — but compressed into weeks, and directed toward a specific functional goal.

Introduce random mutations

Error-prone PCR (epPCR) amplifies the target gene with a deliberately elevated error rate, creating a library of sequence variants — each a slightly different version of the protein. In principle this can generate libraries of thousands of variants; what this project actually recovered was far smaller — see Results for the real number and why. This is analogous to natural mutation, but we control the rate.

Screen colonies on indicator plates

Every colony in the library — not just survivors of a growth challenge, an approach we tested first and found too weak to enrich effectively — is spread onto plates carrying a colour-change substrate: Congo red for retained endoglucanase activity, X-gluc for the β-glucosidase-like activity being evolved. This narrows the library down to a shortlist worth quantifying.

Quantify the shortlist

Shortlisted colonies are grown in liquid culture and lysed, then assayed against a fluorogenic substrate (MUC) to put an actual number on activity rather than relying on a plate colour by eye. This is the step that turns "looked darker on the plate" into a comparable signal.

Sequence and verify

Hits are sequenced to identify which mutations they actually carry. Because epPCR mutagenesis is random, not targeted, sequencing is essential — a colony can look promising on a plate for reasons that have nothing to do with the mutation dry-lab modelling would have predicted.

Why directed evolution works

Natural evolution explores protein sequence space over millions of generations and years. Directed evolution recapitulates this process in the lab by: deliberately narrowing the sequence space explored (we start from a protein that already has related structure), applying a reporter-based screen to find rare interesting variants directly rather than waiting for natural selection, and iterating rapidly (weeks, not millennia).

The key insight is that you don't need to rationally design every mutation. The screening step acts as the filter — you just need to generate enough sequence diversity that beneficial mutations exist somewhere in the library, and then find them.

This project also uses computational structural analysis (AlphaFold3 + docking + FoldX) to identify which residues are worth paying attention to — but epPCR itself mutates the whole gene at random. The computational work can't steer where a mutation lands; it can only flag candidate positions (like Cex N44W or CenA K671S) worth prioritising when deciding what to sequence and follow up on afterward.

Important limitation

Computational predictions guide analysis but cannot replace experimental testing, and in this project they did not dictate which mutations the wet-lab library actually carried — epPCR is random, not targeted. Docking scores predict approximate binding geometry; they do not predict whether a mutation will actually confer catalytic activity in a living cell. See Results for how the dry-lab predictions and the actual wet-lab mutations compared.

Scientific foundation

Key literature this project builds on

2016

Duedu & French — building CxnA

The original construction of CxnA: fusing the Cex (GH10) and CenA (GH6) catalytic domains from Cellulomonas fimi into one engineered bifunctional polypeptide, and providing the sequence and expression system this project uses as its starting point.

Role in this project: Provides the wild-type CxnA baseline — gene sequence, expression protocol, confirmed activities, and domain architecture.
1998/2000

Notenboom et al. — GH10/Cex crystal structure

Crystal structure determination of the Cex exoglucanase domain (PDB: 1FHD and related entries), establishing the active-site architecture including the catalytic glutamates. [3]

Role in this project: Provides the gold-standard crystal structure used to validate AlphaFold3 predictions (RMSD 0.293 Å). Establishes GLU127 and GLU233 as the catalytic dyad.
1996 / 2010

Damude et al.; Cockburn et al. — CenA/GH6 mechanism

Characterisation of GH6-family endoglucanase catalysis, including the catalytic residue architecture used to annotate CenA's active site. [4][5]

Role in this project: Identifies CenA's catalytic residues (Asp595, Asp631, Asp666, Asp771) used in the dry-lab structural analysis for Project B.

In-text citations on this page jump to the summaries above. The complete bibliography, including references not discussed in depth here, is on the Results page.