Methods
Dry lab & wet lab methodology
What we actually ran, as reported in the project's final write-up — not an idealised plan. Every method links to where its results live on the Results page.
Project workflow
One mutant library, three screens, then sequencing
Dry-lab structure prediction and docking ran as its own track to prioritise where mutations might matter — but the actual mutagenesis was random (error-prone PCR), not computationally targeted. A single mutant library was screened through Congo red, X-gluc, and the MUC assay in sequence, each step narrowing down which colonies were worth sequencing.
Computational methods
Dry lab: structure prediction, docking, and mutation screening
Structure prediction (ColabFold)
Full-length CxnA was predicted with ColabFold (AlphaFold2 backend). The top-ranked model (rank_001) was coloured by domain — Cex, the carbohydrate-binding module, and CenA — and rendered in ChimeraX.
Output quality
- pLDDT: Cex 89.30, CenA 93.50 — both high confidence within each domain's own fold
- Global pTM: 0.5 — low, reflecting genuine flexibility in the linker rather than a modelling failure
- PAE (predicted aligned error): mean 26.63 Å, median 27.20 Å between domains
What this does and doesn't tell us
The internal fold of each domain can be trusted individually. The relative orientation and position of Cex versus CenA in the full-length model should not be read as an accurate picture of CxnA's true 3D architecture — only the per-domain structures were used for the docking work below.
Validating the Cex model against a real crystal structure
The predicted Cex domain was superimposed onto the published crystal structure of
Cex in complex with a xylobiose-derived imidazole inhibitor
(PDB: 1FHD, Notenboom et al., 1998/2000).
- Tool: PyMOL
align(iterative outlier rejection) - Result: RMSD = 0.293 Å over 1,958 matched atoms — the converged, best-fitting core, not a pure Cα figure (1,958 exceeds this domain's actual Cα count)
Interpretation
A separate, full-domain Cα-only superposition, run live in the 3D viewer (see Results, DL-R2), gives 0.579 Å over 311 pairs. The two numbers measure different things — a refined best-fitting core versus the complete backbone — so they aren't in conflict, but they also aren't directly comparable to each other, or to the crystallographers' own crystal-form comparison, which used yet another alignment definition. Either figure supports using the predicted model as a docking receptor in DL-3.
Molecular docking (AutoDock Vina)
Two separate docking studies, one per domain, both using AutoDock Vina (Eberhardt et al., 2021).
- Cex: cellobiose docked into the validated Cex active site, to test whether it can engage a β-glucosidase-type substrate at all. Xylobiose (Cex's natural substrate) was also docked as a comparator.
- CenA: cellotetraose (a four-sugar cellulose fragment) docked into the CenA active site. The ligand was sourced as a 3D SDF file from PubChem (CID 170125) before docking.
Critical caveat: absolute docking scores
Vina's binding free energy scores are semi-empirical and not reliable as absolute predictions. The experimentally measured ΔG for xylobiose in Cex (its real, published preferred substrate) is only −3.29 kcal/mol — yet Vina predicts −6.6 kcal/mol for cellobiose in the same enzyme. Scores here are used only for within-receptor pose ranking, never to claim one substrate binds better than another in absolute terms.
PLIP: protein–ligand interaction analysis
Top-ranked docking poses were analysed with PLIP (Protein–Ligand Interaction Profiler, Adasme et al., 2021) to identify hydrogen bonds, hydrophobic contacts, and other interactions between each ligand and its receptor.
| Domain | Residue | Role |
|---|---|---|
| Cex | GLU127 | Acid/base catalyst — contacted cellobiose in the top pose |
| Cex | GLU233 | Catalytic nucleophile — contacted cellobiose in the top pose |
| CenA | ASP595 | General acid catalyst |
| CenA | ASP631 | Modulates the acid catalyst's protonation state |
| CenA | ASP771 | General base catalyst |
PLIP flagged minor structural issues during analysis on both models; these were confirmed not to affect ligand placement as long as the downloaded structure and ligand coordinates were left unchanged.
Conservation analysis and computational mutation screening
Conservation (both domains)
- Cex: BlastP against 15 homologous GH10 sequences, aligned with a Clustal W-based script; conservation mapped with ConSurf.
- CenA: NCBI BLAST against 41 homologous GH6 sequences (65–95% identity), aligned with Clustal Omega; conservation scored 1–11 with Jalview and mapped with ConSurf.
Candidate mutation selection
- Cex: HotSpot Wizard identified mutable residues near the active site (non-essential, and located in the substrate tunnel and/or catalytic pocket). Candidates were screened for predicted stability with FoldX (5 repair runs per mutation) and re-docked with cellobiose to check binding-score impact. Top candidate: N44W (ΔΔG −3.2 kcal/mol).
- CenA: Candidates were screened by combining FoldX stability (ΔΔG) with re-docking against cellotetraose (ΔVina score), plus a check for whether the catalytic residue geometry held up across repeated modelling runs (Pass / Partial / Fail). Top candidate: K671S (best ΔVina improvement, −0.85 kcal/mol, while passing the geometry check).
Molecular biology
Wet lab: construction, mutagenesis, and screening
Liquid cellobiose growth selection (tested, not used downstream)
Before committing to colony-by-colony screening, we tested whether growth on cellobiose as the sole carbon source could itself serve as a selection — pool a mutant library and let only variants with real β-glucosidase-like activity grow, rather than screening colonies individually.
- Strains: E. coli TOP10 expressing wild-type CxnA, or a secreted β-glucosidase as a positive-control comparator
- Media: M9 minimal medium + chloramphenicol + 0.1% casamino acids, with 0.2% glucose or 0.2% cellobiose as the carbon source
- Conditions: inoculated at OD₆₀₀ = 0.05, grown at 37 °C, 200 rpm
- Readouts: OD₆₀₀ at 24 h; viable counts by serial dilution plating on LB + chloramphenicol
Overnight culture and plate preparation
- Host strain: E. coli TOP10, used for CxnA, His-CxnA (frameshift control), and a known β-glucosidase (positive control)
- Liquid culture: standard LB medium, with or without 0.38 mM IPTG and 40 µg/mL chloramphenicol as needed, 255 rpm, 37 °C
- Plates: autoclaved LB agar microwaved 3 min to re-liquify, cooled, then 40 µg/mL chloramphenicol added before pouring into 9 cm dishes
Plasmid extraction and DNA quantification
- Kit: QIAprep Spin Miniprep Kit (Qiagen), standard centrifuge protocol
- Adaptation: where two pellets needed combining, each was resuspended in 125 µL Buffer P1 and the two suspensions transferred into the same microcentrifuge tube before continuing
- Quantification: Nanodrop, with Buffer EB as the blank
An unplanned finding
Gel electrophoresis and sequencing revealed an extra ~1 kb insert in the pSB1C3 backbone. BLASTp matched it to a duplicated chloramphenicol-resistance pseudogene — see Results for the plasmid map and discussion.
Mutant library generation — epPCR and Gibson Assembly
High-fidelity backbone
- Primers Start REV 2 and End FOR 2, Q5 High-Fidelity DNA Polymerase with GC Enhancer (NEB #M0491), annealing at 62 °C
- 50 µL product + 6X Purple Loading Dye (Roche), run on 0.8% agarose gel alongside a 1 kb ladder, visualised under UV
- Gel-purified with the Wizard SV Gel and PCR Clean-Up System (Promega); final elution incubation extended to 5 min
Codon-optimised insert
- Ordered as two synthetic gBlocks (OpPart1_CxnA, OpPart2_CxnA) to work around CxnA's high GC content (~71%)
- 2 ng of each gBlock + 71.5 ng Q5-amplified backbone, joined with the NEBuilder HiFi DNA Assembly Reaction Protocol
Error-prone PCR
- Taq DNA Polymerase Standard Protocol (Roche #11147633103), primers Start FOR Hx2 and End REV 1
- 5% DMSO, 0.32 mM MgCl₂, MnCl₂ gradient 0–0.1 mM to tune mutation rate, annealing at 57 °C
- Products gel-excised, quantified, then ligated to the backbone with a Gibson Assembly kit (NEB)
Honest limitation
Both resulting libraries were smaller than planned — 24 colonies (codon-optimised) and 15 colonies (native sequence) — due to the GC-content amplification difficulty and several troubleshooting delays. See Results for what this means for the statistics downstream.
CMC plates and the Congo red assay (endoglucanase screen)
- Plates: LB agar + 0.2% w/v carboxymethylcellulose (CMC) + 0.38 mM IPTG + 40 µg/mL chloramphenicol, sectioned into four quadrants
- Inoculation: both colony-picking directly onto plates and spotting 1 µL of liquid culture were tested
- Incubation time: a timed 24 h experiment established that 12 h at 37 °C was sufficient before staining
- Staining: flood with 0.1% Congo Red for 15 min, destain with 1 M NaCl for 15 min, measure zones of clearance
X-gluc screen (β-glucosidase activity)
- Substrate: 5-bromo-4-chloro-3-indolyl-β-D-glucopyranoside (X-gluc, Merck), dissolved to 40 mg/mL in 100% DMSO, then diluted 20 µL stock + 80 µL sterile water
- Plates: LB + 40 µg/mL chloramphenicol + 0.38 mM IPTG
- Inoculation: 50 µL spread directly, or 1 µL drops across four quadrants
- Development: overnight growth at 37 °C, then diluted X-gluc either spread over the whole plate or dropped (5 µL) onto individual colonies
MUC assay (exoglucanase quantification) and the MUG assay attempt
Lysate preparation
- Overnight cultures grown with 0.38 mM IPTG from inoculation
- Harvested by centrifugation, washed once in PBS, resuspended in BugBuster lysis reagent + Roche cOmplete EDTA-free protease inhibitor
- Normalised to 5.56 OD₆₀₀·mL of culture per mL of lysis reagent, so every lysate receives a matched cell input
- Lysed 30 min at room temperature with gentle agitation, clarified by centrifugation (4000g, 5 min), held on ice; all compared lysates were prepared and assayed the same day
- Total protein estimated by A280 on Nanodrop (used only as a rough check — see the honest note on this below)
MUC reaction
- Reaction: 200 µL total = 10 µL lysate + 188 µL PBS + 2 µL MUC reagent (10 mM stock in DMSO) → 200 µM MUC in-well
- Plate: Greiner 96 F-bottom
- Reader: BMG Labtech FLUOstar Omega, bottom optic, Ex 310-10 nm / Em 460 nm, 37 °C, read every 2 min for 198 min
- Gain: 500 for the screening plate; 1000 for earlier method-development plates
- No excitation filter closer to the 4-MU anion's true absorbance maximum (≈360 nm) was available on the instrument, so all readings used 310 nm
Calibration and controls
- 4-MU standards (0–6.25 µM in-well) prepared by serial dilution in DMSO, then transferred into the assay matrix — included on every plate in both PBS and frameshift-lysate backgrounds
- Enzyme-amount proportionality established by mixing active lysate with frameshifted (CxnA-His) lysate at fixed ratios (1:0, 1:1, 1:3, 1:7, 1:15) while holding total lysate at 10 µL — a volume-varying alternative was tested first and rejected (see Results)
The MUG assay was attempted, using the wrong substrate
We intended to run this as a MUG assay for β-glucosidase activity, but the bottle on the shelf turned out to be the wrong substrate — 4-Methylumbelliferyl-β-D-galactopyranoside (MUGal, a β-galactosidase substrate) instead of MUG. It was assayed exactly as above, at 200 µM in-well, with one plate adding substrate/lysate dilution series and boiled, −IPTG, and cell-free-medium controls once the signal looked wrong and needed explaining. The reaction completed before the first read, so plateau fluorescence (not initial rate) was reported. The signal traced to host β-galactosidase activity via lacZα α-complementation in TOP10, not to CxnA — full story in the Results write-up and the 15 September notebook entry. The MUC results are unaffected: the frameshift control sits at under 1% signal on MUC, so β-galactosidase activity doesn't touch the cellobioside substrate.
Sequencing and mutant validation
Colonies flagged by the Congo red, X-gluc, and MUC screens as interesting — retained activity, lost activity, or an above-baseline MUC signal — were grown in liquid culture for plasmid isolation.
- Plasmid isolation: QIAprep Spin Miniprep Kit (Qiagen)
- DNA quality check: Nanodrop 2000 (Thermo Fisher) — concentration, A260/A280, and A260/A230 ratios, before committing samples to sequencing
- Sequencing: Sanger sequencing via MRC PPU DNA Sequencing and Services