Cambridge · iDEC 2026

Flexible proteins.
New possibilities.

Exploring disordered protein tags through computational design and experimental selection.

Can intrinsically disordered protein tags help proteins remain soluble, or help a cell withstand freezing? We explored these questions by designing sequence libraries, applying selection and tracking which variants became enriched.

Cambridge iDEC 2026: Darwin, Cambridge architecture and an evolution-inspired protein illustration
4 design branchesNEXT · SSB · De novo · CAHS1
2 selection strategiesAntibiotic exposure · Freeze–thaw
1 shared approachDesign · Build · Select · Analyse

01 / The problem

Producing a protein is only the beginning.

Recombinant proteins are valuable tools for research and biotechnology, but expression does not guarantee a useful product. Proteins can aggregate, become insoluble or lose activity. Finding conditions that preserve their function can require substantial optimisation.

Fusion tags offer one possible route. Our project asks whether intrinsically disordered proteins and protein regions can provide useful starting points for developing tags that support protein function or cellular stress tolerance.

02 / The design

Explore the sequence through directed evolution.

Intrinsically disordered proteins do not adopt a single stable three-dimensional structure under the conditions in which they are disordered. Their flexibility and sequence composition make them candidates for several protective roles. For solubility tags, one proposed mechanism is an “entropic bristle”: a mobile peptide that can hinder aggregation-promoting contacts around its fusion partner.

We combined natural templates with a motif-based de novo design. Computational tools generated amino acid variants and filtered them using descriptors such as hydrophobicity, estimated charge and secondary-structure propensity. Degenerate DNA encoding provided a practical way to synthesise libraries, while also introducing combinations beyond the original filtered catalogue.

01

Natural peptide extension

NEXT

A 52-residue construct derived from the N-terminal extension of a bacterial carbonic anhydrase. We explored variants of this solubility-tag template through β-lactamase selection.

Explore the NEXT design
02

Disordered bacterial tail

SSB

A 42-residue C-terminal fragment of E. coli single-stranded DNA-binding protein. We used its flexible tail as a starting point for a fusion-tag library.

Explore the SSB design
03

Motif-based sequence design

De novo

New candidate sequences assembled from charged, polar and proline-containing motifs. A culture identity mix-up prevented experimental conclusions about this branch.

Explore the de novo design
04

Tardigrade-inspired protection

CAHS1

A separate branch focused on cellular freeze–thaw tolerance. We generated CAHS1-derived variants and examined their representation across selection rounds.

Explore the CAHS1 design

03 / The selection

Two selection strategies.

NEXT + SSB

Function under antibiotic exposure

Candidate tags were fused to aggregation-prone L76N TEM-1 β-lactamase, an enzyme that can provide resistance to carbenicillin. We monitored culture turbidity and compared sequence representation in cultures grown with and without carbenicillin.

The aim was to enrich variants associated with function in this system. Growth and sequencing provide an indirect selection readout; they do not by themselves measure soluble enzyme yield.

CAHS1

Recovery after freeze–thaw stress

Freeze-thaw stress increases the level of aggregation of certain proteins, such as LDH. Cells expressing CAHS1-derived variants underwent freeze–thaw selection. Viable-cell recovery and sequencing were used to investigate survival and changes in library composition across rounds.

A preliminary comparison of the original CAHS1 sequence with empty pBAD provided evidence that the system could distinguish different survival outcomes.

From sequence design to candidate identification Four stages: computational design, library construction, experimental selection, and sequencing analysis. A connecting line draws as the diagram enters view. 1 2 3 4 DesignExpressSelectAnalyse
Sequence design and experimental selection generate candidates for subsequent individual validation.
Read the experimental design

04 / The evidence

Selection changed which sequences were represented.

We addressed two questions: did the overall sequence composition change, and which groups of similar sequences became more or less abundant?

Compare composition

We compared amino acid variant counts using Pearson’s χ² statistic, with fixed-margin Monte Carlo simulations to assess the discrepancy under an equal-composition model when expected counts were sparse.

Explore sequence properties

Within each library, we clustered the pooled untreated and selected sequences once, using a shared set of scaled descriptors. PCA provided a two-dimensional view, while read counts measured the abundance of each fixed property group.

NEXT

A property group became enriched

The C2 group increased from 6.1% to 29.6% of reads. NEXT-072 was prioritised as an individual candidate based on its increase in relative representation.

SSB

One variant dominated the selected sample

SSB-054 accounted for 72.6% of selected reads. Its dominance strongly influenced the read-weighted property summaries.

CAHS1

Sequence change and survival evidence

Library sequencing showed changes across selection rounds. In a separate preliminary deep-freeze comparison, cells expressing original CAHS1 survived better than the empty-vector control.

These are exploratory findings from single biological replicates. Enrichment identifies candidates for follow-up; it does not establish reproducible improvement in protein solubility or isolate the mechanism responsible.

Explore the results

05 / What comes next

From enriched candidates to tested functions.

The next step is to reconstruct prioritised variants and test them individually alongside parental sequences and appropriate controls. Replicated assays of soluble protein yield, enzyme activity or freeze–thaw survival would help establish whether enrichment corresponds to the intended phenotype. Testing additional fusion partners would then address how transferable a useful tag might be.

Our contribution is a documented workflow linking computational library design, experimental selection and sequence-level analysis, as well as a set of candidates and questions for future work.