A phosphorylation switch in PAGE4 drives MED12-mutant fibroid pathogenesis
This article has been Reviewed by the following groups
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
- Evaluated articles (Review Commons)
Abstract
Recurrent somatic mutations in MED12, found in ∼70% of uterine leiomyomas (ULs), define the dominant molecular subtype of these highly prevalent tumors, yet the downstream effector mechanisms remain poorly understood.
Using an integrated multi-omics workflow, encompassing discovery-phase DDA proteomics of matched leiomyoma–myometrium pairs, validation-phase DIA proteomics across genetically stratified cohorts (MED12 p.G44D, RAD51B–HMGA2, IRS4/FH subtypes), phosphoproteomics, immunohistochemistry, and AP-MS interactomics, we identify prostate-associated gene 4 (PAGE4) as a central effector of MED12-mutant UL pathogenesis.
PAGE4 emerged as one of the most significantly upregulated proteins in MED12-mutant ULs and harbored the highest-occupancy hyperphosphorylation sites (T51, T85) in the tumor phosphoproteome, a pattern confirmed by immunohistochemistry and orthogonal phosphoproteomics. Kinase-substrate enrichment nominated HIPK2 as the primary upstream kinase, with supporting evidence from the broader CMGC family. Subsequent analysis revealed that phosphorylation acts as a molecular switch, fundamentally restructuring the PAGE4 interactome to favor the Mediator complex and RNA Pol II transcriptional machinery. This rewiring was functionally validated by ChIP-seq and luciferase reporter assays, which demonstrated corresponding shifts in transcription factor occupancy and downstream pathway activity.
Collectively, these data establish phospho-PAGE4 as a critical mechanistic node downstream of MED12 mutation, expand PAGE4 biology from prostate cancer to female reproductive tumors, and nominate it as a mechanistic biomarker and candidate therapeutic vulnerability in the most prevalent molecular subtype of uterine leiomyomas.
Article activity feed
-
Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.
Learn more at Review Commons
Reply to the reviewers
Response to Reviewers # # Reviewer #1
We thank Reviewer #1 for the positive assessment, in particular for recognizing that the sex-independent role of PAGE4 is clearly supported by multiple layers of protein evidence, that the proteomic, phosphoproteomic and interactomic techniques are conducted at an expert level, and that the reporting of results is appropriate. We have addressed the constructive points on outlier testing and methods completeness in full.
__Reviewer comment: __Figure 1B: the first PCA component is dominated by a single sample far to the right (PC1 33.9%). Was an outlier test (Mahalanobis distance, Z-score on PC1) performed? If so it …
Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.
Learn more at Review Commons
Reply to the reviewers
Response to Reviewers # # Reviewer #1
We thank Reviewer #1 for the positive assessment, in particular for recognizing that the sex-independent role of PAGE4 is clearly supported by multiple layers of protein evidence, that the proteomic, phosphoproteomic and interactomic techniques are conducted at an expert level, and that the reporting of results is appropriate. We have addressed the constructive points on outlier testing and methods completeness in full.
__Reviewer comment: __Figure 1B: the first PCA component is dominated by a single sample far to the right (PC1 33.9%). Was an outlier test (Mahalanobis distance, Z-score on PC1) performed? If so it should be reported. No PCA or outlier test is reported for the phosphoproteomics. The full-proteome and phosphoproteomics data need more analysis detail to ensure reproducibility.
__Response: __We agree, and we have addressed this in two ways. First, we have replaced the Figure 1B proteome PCA with multidimensional scaling (MDS) computed from the top 500 most variable proteins, which displays the paired tumor-myometrium structure directly and is not dominated by a single component; the inputs are stated in the Methods and the Figure 1B legend: MDS was computed from the top 500 most variable proteins using log2-transformed intensities, with no imputation of missing values. Second, to address the reviewer's underlying concern about undetected technical outliers, we inspected the per-sample distribution of log2 protein intensities for every sample in the DIA cohort. One HMGA2-subtype tumor measurement was clearly aberrant at the technical level: its distribution is bimodal and its median is collapsed (~2.9 versus a cohort median of 7.8), the signature of a sample-specific quantification failure rather than biological variation (rebuttal figure below, panel A), whereas the same sample is unremarkable in the phosphoproteome (panel B), showing that the defect is specific to that one proteome measurement and not a property of the tumor. That measurement is not part of the dataset reported in the manuscript: the DIA cohort as analyzed comprises 28 matched tumor–myometrium pairs (56 samples), which is what the Methods, the figure legends and Supplementary Dataset 2 describe throughout. Every remaining sample passes this inspection, so no result reported here rests on a sample with a failed proteome measurement. We provide the per-sample distributions here for the reviewer's inspection rather than in the manuscript, since they document the composition of the analyzed cohort rather than adding to its findings.
Rebuttal figure1* for reviewer inspection only. Per-sample intensity distributions across the DIA cohort. (A) Proteome; (B) phosphoproteome. Each box is one sample's distribution of log2 intensities; the blue dashed line marks the cohort median (7.8). The aberrant sample (red) is a clear technical outlier in the proteome but unremarkable in the phosphoproteome; it is not part of the 28-pair dataset reported in the manuscript.*
__Reviewer comment: __Figure 3A: perform an outlier test for HMGA2 (proteome, left) and for the FH_UL and COL4A5-COL4A6_UL points (phosphosites) that are visually separated. As shown, the PCAs are hard to interpret because points get squeezed onto one axis. Outliers may affect statistics and the conclusions about within-group variation; this should be tested if discussed.
Response: We have addressed both points. The same per-sample inspection described above covers every sample contributing to the Figure 3A proteome, phosphoproteome and phosphosite analyses, and all 28 pairs shown there pass it. The COL4A5-COL4A6 UL and FH UL samples highlighted by the reviewer are technically sound on these metrics and are retained; the text now attributes their separation to inter-tumor heterogeneity in phosphorylation state on the basis of these metrics rather than by assertion, which resolves the biological-versus-technical ambiguity the reviewer identified. Regarding axis compression, we agree that the phosphosite panel is stretched by a small number of samples with genuinely divergent phosphorylation profiles. Because these samples are technically sound, we have retained them rather than rescaling or trimming the axes, which would misrepresent the spread of the data; the proportion of variance explained by each component is stated on the axes so that the scale of the separation can be judged directly. The conclusions drawn from Figure 3 concern the tumor-myometrium contrast within each subtype, which is unaffected by the position of these individual samples.
__Reviewer comment: __SNRNP70_S226 is described as one of the most significantly hyperphosphorylated sites, but Table S1 lists log2FC 5.92, p 0.02, classified non-significant; MARCKSL1 (log2FC 2.14, p 0.016) and NDRG2 are likewise non-significant. The volcano/caption indicate a 0.05 cutoff but Table S1 apparently uses 0.01. Please clarify.
Response: We thank the reviewer for catching this, and we apologize for the inconsistency. On re-checking the source data we confirmed that the significance-classification column in the earlier version of the supplementary table had been generated with a stricter cutoff than the one used for the volcano plots and stated in the captions. All of the sites raised by the reviewer meet the stated criteria (|log2 fold-change| > 2 and p ##
__Reviewer comment: __Add citations for Protein Atlas and ProteomicsDB.
__Response: __Added: Human Protein Atlas (Uhlén et al., Science 2015; proteinatlas.org) and ProteomicsDB (Schmidt et al., Nucleic Acids Res 2018).
__Reviewer comment: __PAGE4 expression is selectively enriched in MED12-mutant UL (page 14): "PAGE4 expression indeed was significantly elevated in MED12-mutant ULs ... ". The figure caption, methods and text do not provide any information of the statistical test used regarding the statement of significant upregulation. Please provide the data or change wording as a statistical significance is implied.
Response: The panel reproduces published RNA-seq data (Berta et al., 2021) descriptively, and the source data do not provide a formal cross-subtype test that we can report. We have therefore softened the wording from “significantly elevated” to “consistently higher” in the Results, and the caption now states that the panel is descriptive and that no statistical test was applied.
__Reviewer comment: __Method section "Liquid chromatography-mass spectrometry (LC-MS)" report the UniProtKB database download date to ensure reproducibility. Does the database contain isoforms and TrEMBL sequences or uses a filtered version (e.g.: status reviewed, canonical?) (report it like on p.36 for the AP-MS/BioID samples). Further add a statement if missing values were imputed on which level (peptide, protein), which normalization strategy was employed, and was there filtering of incomplete values performed. All these analysis steps are common practice for DDA measurements of clinical samples and may affect subsequent statistical finding and thus should be reported in the method section if they were or were not performed (for the PPI data this was reported on p.37, but for the full proteome and phosphoproteome the information provided is not sufficient to judge the validity of data processing). Method section "GO enrichment and statistical analysis" (p.37): Please add citations for the used tools SAINT, CRAPome and Enrichr.
__Response: __We have expanded the discovery DDA proteome and phosphoproteome Methods to the same level of detail as the AP-MS/BioID section: the UniProtKB human database version, download date and entry count; whether isoforms/TrEMBL were included or a reviewed-canonical subset was used; the valid-value filtering rule per group; the normalization strategy; and the imputation method and level. The corresponding DIA workflow is now also described (see Reviewer 2-J). Also added references for SAINTexpress (Teo et al., J Proteomics 2014), CRAPome (Mellacheruvu et al., Nat Methods 2013) and Enrichr (Kuleshov et al., Nucleic Acids Res 2016).
__Reviewer comment: __The Figure 1c: The scale of the (Gene) count does not correspond with the circle sizes of the plot. It is thus not possible to judge the number of genes per category. This should be sized equally as else it is not possible for a reader to judge the number of genes per enriched terms.
__Response: __We have corrected the GO enrichment panel (Figure 1D) so that dot area is drawn to a single consistent scale with an accurate, labeled size legend, allowing the gene count per term to be read directly.
__Reviewer comment: __Figure 1G: Expression levels for ProteomicsDB. As the log2 of the expression level is used, it is not clear if the white squares correspond to missing values in the ProteomicsDB, or if the value was very low. Maybe indicating absent values in a grey color scheme makes this more explicit.
__Response: __We have clarified this in the caption rather than by recolouring, because in this figure white already carries a defined meaning in both panels: in the Human Protein Atlas panel it is the explicit “not detected” category of the key, and in the ProteomicsDB panel the log2 intensity scale begins at zero, so a white cell corresponds to a value of zero, that is, no detected expression, rather than to an absent measurement. The caption now states this for both panels, so a white square can no longer be confused with a very low value.
Reviewer comment:* Figure 1C: the caption indicates a "gray box" but it seems that PAGE4 was highlighted as a red dot, and a white box was added for indicating the gene name.*
__Response: __The caption has been corrected to match the figure (PAGE4 shown as a highlighted point with a labeled callout), and color terminology has been harmonized across all captions.
__Reviewer comment: __Figure 3D is described as a volcano plot but is a horizontal scatter with a significance threshold rather than a classic volcano (which would have adjusted p-value on the axis).
Response: The reviewer is correct. These panels display log2 fold-change per subtype with significance indicated by color and do not plot a p-value axis, so “volcano plot” was the wrong term. We have relabeled them as grouped dot plots in the figure legend and at the two places they are referred to in the Results. Figure 2C and Figure 5A, which are conventional volcano plots with a -log10 p-value axis, retain that description.
__Reviewer comment: __The PPI network has too many edges and nodes to read. It could help to collapse categories/proteins or provide a subnetwork, with the full network in the supplement.
__Response: __We have moved the dense, full node-level PAGE4-centered network out of the main figure and into the supplement, where it now appears as Figure S3A. The main figure retains a complex-level view of these interactions as a focused interaction matrix (revised Figure 4E), in which preys are grouped by annotated functional complex with confidence-weighted encodings. This is the same change requested by Reviewer 3 (point G).
Reviewer #2
We thank Reviewer #2 for the positive overall assessment, that we convincingly identified a novel role for PAGE4 in uterine leiomyoma and Mediator transcriptional processes, that this represents an important contribution to understanding UL pathobiology, and that the dataset is very comprehensive. We are grateful for the detailed, panel-by-panel reading and the constructive suggestions on figure design and methods organization, all of which we have adopted.
__Reviewer comment: __The validation cohort is not technically a validation cohort, it uses the same 9 MED12-mutant UL samples and so cannot validate PAGE4 upregulation/phosphorylation in an independent cohort. A second MED12-mutant UL cohort should be used.
__Response: __We thank the reviewer for raising this, and we apologize that our original Methods wording was ambiguous. The discovery DDA, and validation DIA analyses were in fact performed on independent sets of MED12-mutant UL/myometrium samples (n = 9 each), with no patient shared between the sets; we originally sized (MED12-mutant) the DDA/IHC and DIA sets to match the nine discovery pairs, and then also had the opportunity to analyse also the other mutations causing UL. The DIA experiment therefore does provide independent validation of PAGE4 upregulation and Thr51/Thr85 phosphorylation in a separate cohort of MED12-mutant tumors, while additionally establishing subtype specificity against the HMGA2, FH and COL4A5-COL4A6 subclasses. We have revised the Methods to state the cohort composition explicitly so that the independence of the discovery (DDA) and validation (DIA) MED12 sample sets is unambiguous, and independent confirmation is further supported by the external RNA-seq dataset (Berta et al. 2021) and the IHC validation.
__Reviewer comment: __No meaningful change in peak width or height is readily observable. Statistical tests would be required to make the claims believable.
__Response: __We agree, and we have replaced visual assertions with quantification and statistics, reporting effect sizes alongside p-values. We now distinguish signal amplitude (which changes) from peak shape (which does not).
Amplitude is reduced (supported). Across a common set of 19,942 GENCODE v38 protein-coding TSSs, median promoter-proximal RNAP II signal (TSS ±1 kb) decreased by 10.4% in MED12 G44D, 16.5% in PAGE4 S9D/T51E/T85E and 21.2% in PAGE4 S9A/T51E/T85A relative to the matched wild type; the area under the curve decreased by the same margins, and median peak height (TSS ±250 bp) decreased by 4.4%, 20.4% and 19.3% respectively. All reductions are statistically significant (two-sided Wilcoxon rank-sum test, Benjamini–Hochberg-adjusted p Shape is unchanged (we concede this). Median TSS-derived FWHM was 201 bp in all five conditions (Δ = 0); the very small FWHM p-values arise from the large number of regions tested (n ≈ 17,000) and do not indicate a biologically meaningful shape change. MACS2 peak-width medians shifted inconsistently (+23 bp for G44D, +1 bp for S9D, −49 bp for S9A), i.e. no consistent broadening or narrowing.
We removed the statements that all four metrics concord and that peaks are narrower/broader. The revised Results now state that promoter-proximal RNAP II signal amplitude is modestly reduced, whereas peak shape (FWHM and MACS2 peak width) shows no consistent change, and we report effect sizes (median % change) for every metric. The ChIP-seq analysis pipeline (alignment, MACS2 peak calling, normalization, quantification) has been added to Methods, which previously lacked it.
__Reviewer comment: __HIPK2 was proposed as the candidate kinase (Fig 3E) but there is no HIPK2 on the heatmap; instead CLK2 is in red with an increased score. HIPK2 and CLK2 are different, unrelated kinases. Please rectify.
__Response: __We apologize for the inconsistency, which our cross-audit confirmed. We have replaced the kinase-family enrichment heatmap (original Figure 3E) with a per-site kinase-prediction map for PAGE4 S9, T51 and T85, now shown as the lower panel of Figure 4A, and revised the text so that figure and text agree. Both sites are proline-directed (S/T-P) CMGC substrates; the CLK and HIPK families score highest, with CLK2 and HIPK1 the most probable kinases, as both have been experimentally validated as PAGE4 kinases at these sites (Kulkarni et al., 2017; HIPK1 also reported at T51). We no longer single out HIPK2: motif scoring cannot resolve closely related paralogs, so HIPK2/HIPK3 and other CLKs remain possible but are not named individually, while the prior experimental literature points to HIPK1. The prediction is now explicitly framed as in silico and requiring direct validation (in vitro kinase assays with recombinant CLK2/HIPK1, and inhibitor or genetic perturbation).
__Reviewer comment: __The text states Mut1 retained 73% of PAGE4-WT prey proteins, but Fig 4B shows Mut1 with 25 preys vs 149 for WT, 25/149 = 17% retained, an 83% decrease. Please rectify.
__Response: __The reviewer is correct, and the original 73% was an error. We recomputed AP-only interactor retention directly from the final filtered interactor table (Supplementary Dataset 3). Defining retention as the fraction of PAGE4-WT AP high-confidence interactors (gene-level, n = 149) that are also recovered for Mut1, only 20 are shared, a retention of 13% and a net loss of 87% of stable associations; Mut1 captured 25 AP high-confidence interactors in total (bar plot, revised Figure 4C). The text and the Figure 4C legend now report these values consistently, and the statement about which categories are preferentially lost versus retained is tied to Supplementary Dataset 3 rather than asserted.
__Reviewer comment: __The text states the MED12-G44D bait is strongly reduced vs MED12 WT (Fig 5A), but the Fig 5A right panel does not support this. Please rectify.
__Response: __We thank the reviewer, on checking the data, the original statement was in fact wrong in the opposite direction, and we have corrected the text. Quantifying MED12 bait recovery directly from Table S3, the MED12-G44D bait is recovered at levels equal to or higher than MED12-WT (average spectral counts: AP-MS 519 vs 339; BioID 481 vs 191), not reduced. Despite this comparable-to-higher bait abundance, MED12-G44D captures markedly fewer high-confidence interactors (AP-MS 38 vs 115; BioID 297 vs 362). The Results now state this explicitly: the G44D mutation reduces MED12's high-confidence interaction repertoire while the bait itself is well expressed, so the loss of interactors is a genuine effect rather than a bait-recovery artifact. Figure 5A is a differential-interactor volcano (mutant vs WT) in which the bait's own abundance is not directly legible, which is why it appeared not to support the now-removed reduction claim. Bait spectral counts for every bait and both acquisition modes are reported in Supplementary Dataset 3, so bait recovery can be checked directly for each construct. In the interest of full transparency we note that the same check gives a different answer for two of the PAGE4 constructs: Mut1 and Mut6 are recovered at roughly ten-fold lower spectral counts than PAGE4 WT in both modes. We have therefore added an explicit limitation to the Discussion stating that reduced construct abundance or stability may contribute to the loss of stable interactors and to the compartmental shift observed for these two variants, and we have tempered the corresponding claims.
__Reviewer comment: __The text states MUT1 has upregulated base-excision repair and SUMOylation, but the Fig S3 heatmap supports this for MUT6, not MUT1. Please rectify.
__Response: __Thank you. We re-examined the pathway enrichment heatmap, which is Figure S4 in the revised manuscript, and confirm the reviewer is correct: base-excision repair, translesion synthesis and SUMOylation of DNA damage response proteins are enriched for the S9A/T51E/T85A variant (Mut6), not for Mut1. The Results text has been corrected so that text and figure agree.
__Reviewer comment: __Figure 4E shows the full bait-prey graph and becomes a hairball; bait names are illegible, edge widths/colors and AP/PL/Both cannot be distinguished. Suggest aggregating preys by annotated complex with a single summarized weighted edge per bait-complex, and showing only differentially associated preys; provide the full network in the supplement.
__Response: __We agree. The dense, full node-level network that produced the hairball has been moved out of the main figure and is now provided as a supplementary figure (Figure S3A) with a clearly keyed AP/PL/Both encoding. In the main figure, the corresponding interactions are presented at the level of annotated complexes as a focused interaction matrix (revised Figure 4E), which avoids the unreadable node-level graph while preserving the key bait-complex relationships.
__Reviewer comment: __The figure presents PAGE4 phosphosites but is labeled “GAGE,” a different (unused) name; importantly the figure appears copied directly from PhosphoSitePlus without permission. Permission should be obtained.
__Response: __We thank the reviewer and have resolved both issues. First, we replaced the PhosphoSitePlus-derived schematic with an original figure generated by us from the underlying site annotations, so no third-party copyrighted image is reproduced, removing the permissions concern entirely. Second, the “GAGE” label is showing the GAGE domain (PAGE4 belongs to the GAGE/PAGE family); all panel labels now read PAGE4 consistently, with the alias noted once in the text.
__Reviewer comment: __Inferring phosphorylation involvement from phosphomimetic changes is not formally valid and should at least be stated as a hypothesis.
__Response: __We agree and have reframed accordingly. We now state explicitly that phosphomimetic (S/T→D/E) and phosphodead (S/T→A) substitutions approximate the charge state of (de)phosphorylated residues but do not reproduce native, dynamic phosphorylation; the phosphorylation-dependence conclusions are therefore presented as a hypothesis supported by, but not proven by, these surrogates.
__Reviewer comment: __The MS methods are insufficient to reproduce the experiments. DIA methods are missing entirely; with both DDA and DIA used across total, phospho, AP and PL experiments, the methods should be reorganized by experiment type, specifying what was done in each mode. State how total and phospho data were normalized and integrated; how PL data were analyzed and whether identically to AP; whether SAINT used spectral counts or intensities; and, since the baits are nuclear, whether NLS-containing controls were used.
__Response: __We have substantially expanded and reorganized the MS Methods by experiment type and acquisition mode. (1) DIA acquisition and analysis for the validation proteome and phosphoproteome are now fully described, instrument, gradient and window scheme, search/quantification software and settings, library strategy, FDR and normalization. (2) For each dataset we state the acquisition mode: discovery proteome and phosphoproteome = DDA (Q-Exactive); validation proteome and phosphoproteome = DIA-PASEF (timsTOF Pro); AP-MS and PL-MS = DDA-PASEF (timsTOF Pro/Pro 2); host-cell-line proteome = DIA on an Orbitrap Astral, for which the full acquisition and DIA-NN parameters, including database version, entry count and the absence of isoforms, TrEMBL entries and imputation, are now given in Methods. (3) Total- and phospho-proteome integration: phosphosite intensities were normalized to matched parent-protein abundance for the protein-adjusted, site-level testing, and we describe the normalization, batch handling and integration of the two layers. (4) AP vs PL: we state explicitly that AP and PL data were processed through the identical pipeline (QRILC imputation → median normalization → SAINTexpress → CRAPome filtering), noting any differences. (5) SAINT input: we clarify that SAINTexpress was run on spectral counts (and CRAPome filtering used average spectral-count fold-change ≥ 3); the QRILC/normalization step applies to the quantitative interactor comparisons, and the previous ambiguity between counts and intensities is removed. (6) Nuclear-bait controls: GFP-MAC3 served as the negative control, and compartment-matched contaminants (nucleolar/chromatin) are additionally controlled by CRAPome frequency filtering; we discuss this explicitly and note its limitation for nuclear-bait specificity.
__Reviewer comment: __Overall the manuscript has many issues with data presentation. To illustrate just a few:
- Fig. 3B and 3D are described as volcano plots. They are not volcano plots.
- Fig. 3B legend mentions a lower panel. There is no lower panel for Fig. 3B.
- Fig. 3A, please use different shape markers for the UL subtypes and suggest connecting with edges to thus support their text conclusions of separate groupings.
- For Fig. 3E the colors for the families and groups are not distinguishable (and the font is too small). Please rectify.
- Please label the units and scale for Fig. 6A.
- Please mention the MAC3 tag uses Ultra-ID to thus justify the 10 min labeling time.
- In the relevant protein and phosphoproteomic datasets and figures please mention the number of proteins and phosphosites obtained to thus allow an interpretation of the data quality. -Table 4 is in a format that is uninterpretable.
- The text and figures interchangeably use the mutant number (e.g. MUT1) or mutant composition (e.g. S9D_T51E_T85E - for MUT1), this makes reading and interpreting the manuscript very challenging. Suggest using one format.
__Response: __We have addressed each item:
- Fig 3B and 3D “volcano” plots: corrected. Both panels are now labeled grouped dot plots in the figure legend and in the Results, matching our response to Reviewer 1; Figure 2C and Figure 5A, which do plot a -log10 p-value axis, retain the term volcano plot.
- Fig 3B legend “lower panel” with no lower panel: the legend conflated the protein panel (3B) with the phosphosite panel (3D); the caption now matches the actual panels.
- Fig 3A subtype markers: we have retained the colour-coded points. Each of the eight groups already has its own colour and the legend states the group and n for each, and the eight-level palette was chosen so that each subtype's tumour and myometrium share a colour family, which is the comparison the panel is meant to support. We considered adding shape markers and 95% confidence ellipses, but at these group sizes (n = 5 for two subtypes) an ellipse conveys more confidence in the grouping than the data warrant, and connecting edges between unpaired samples would imply a trajectory that does not exist. The panel is therefore presented as a descriptive overview of sample structure; the subtype and tumour-versus-myometrium differences that the manuscript actually concludes on are tested statistically and reported in Figure 3B and Supplementary Dataset 2, not inferred from visual separation in the PCA.
- Fig 3E colors/fonts: the original Figure 3E no longer exists; it has been replaced by the site-level kinase-prediction panel that is now the lower panel of Figure 4A, which uses a larger font and a distinguishable family/group palette.
- Fig 6A units/scale: the Figure 6A legend now specifies that the heatmap shows MACS2 fold-enrichment (FE) signal over local background (macs2 bdgcmp, FE mode), in 50-bp bins across ±5 kb from GENCODE v38 protein-coding TSSs, with the display color scale capped at 20 FE units (values above 20 capped only for visualization; quantitative analyses used uncapped values).
- MAC3 labeling time: we now state that the MAC3 tag incorporates the UltraID biotin ligase, justifying the 10-minute labeling window.
- Dataset sizes: the number of proteins and phosphosites identified in the discovery datasets is stated in the Results, and the complete per-feature quantifications for both the discovery and validation datasets, from which the totals can be read directly, are provided in Supplementary Datasets 1 and 2.
- Table S4 formatting: reformatted into an interpretable layout.
- Mutant nomenclature: standardized to “Mut1 (S9D/T51E/T85E)” on first mention and “Mut1” thereafter. The full Mut1-Mut8 mapping is now stated explicitly in the Results at the point the panel is introduced, and the same convention is applied throughout.
Reviewer #3
We thank Reviewer #3 for the positive and thorough assessment, in particular for recognizing that we convincingly demonstrated the specific elevation of PAGE4 mRNA, protein and T51/T85 phosphorylation in MED12-mutant UL, that this points to a broader, sex-independent function for a traditionally male-specific marker, and that the work will be of significant interest to the UL research community. We address the major and minor points in turn.
__Reviewer comment: __The interactome, reporter and ChIP-seq studies relied exclusively on tagged, exogenously expressed proteins. Endogenous PAGE4/MED12 could substantially affect the observed networks. Measure endogenous expression in the Flp-In 293 T-Rex cells and discuss its influence. Did phosphovariants affect PAGE4 nuclear localization? Comparing T51/T85 phosphorylation between WT and G44D MED12 cells would be informative.
__Response: __We have characterized endogenous PAGE4 and MED12 in the parental Flp-In T-REx 293 line. By single-shot DIA proteomics of the parental cells on an Orbitrap Astral, PAGE4 was below detection, indicating that the tagged PAGE4 constructs are expressed against an effectively null endogenous background, consistent with PAGE4's restricted cancer-testis expression; the reported PAGE4 interactions, localization and reporter effects are therefore not confounded by an endogenous PAGE4 pool. MED12, an essential Mediator subunit, is endogenously expressed, so the tagged MED12-WT and MED12-G44D constructs are present alongside the endogenous protein, as we now note in the Discussion. We have also assessed whether PAGE4 phosphovariants alter nuclear/cytoplasmic distribution using our developed MS microscopy analysis (Liu et al, 2018, Nature Communications), We also agree that comparing PAGE4 T51/T85 phosphorylation between WT and G44D MED12 backgrounds would be informative; this requires phosphosite-specific reagents that we do not currently have, and we have flagged it explicitly in the Discussion as a priority for follow-up rather than claiming it here. The revised Discussion frames the heterologous overexpression system and its bearing on interpretation accordingly.
__Reviewer comment: __Turunen et al. (2014, PMID 24746821) reported mutant MED12 interactomes using a similar strategy. Comparing with this dataset would assess reproducibility and robustness and add confidence.
__Response: __We thank the reviewer for this excellent suggestion, and we note that a co-author of the present study also co-authored Turunen et al. (2014). We have compared our MED12 WT and G44D affinity-purification (AP-MS) interactors with the normalized spectral-count data of Turunen et al. (2014, Table S1), normalizing both datasets to the MED12 bait (see the rebuttal figure below). The two studies are concordant on the central finding: the G44D mutation specifically reduces association of the Mediator kinase (CDK) module while core Mediator is retained. In our data the bait-normalized G44D/WT ratio for CDK8 is 0.45 and for Cyclin C (CCNC) is 0.56 (Turunen et al.: 0.33 and 0.36, respectively), and CDK19 — the module member lost most completely in Turunen et al. (ratio 0.00) — was not detected as a high-confidence interactor of G44D in our AP-MS data. By contrast, core Mediator subunits are retained in both studies (median G44D/WT ratio 0.87 across 23 shared subunits in our data, versus 0.92 in Turunen et al.). Thus our MED12 G44D interactome independently reproduces the selective CDK8/Cyclin C uncoupling reported by Turunen et al. (2014), while our combined AP-MS/PL approach extends this to the PAGE4 variants and to transient/proximal associations. We have added a sentence to the Discussion noting this concordance; the comparison figure is provided here for the reviewer's inspection.
Rebuttal figure 2. Concordance of the MED12 G44D interactome with Turunen et al. (2014). Both datasets are normalized to the MED12 bait and expressed as the G44D/WT ratio. (A) Mediator kinase (CDK) module: CDK8 and Cyclin C (CCNC) are reduced in both studies; CDK19, the module member lost most completely in Turunen et al. (ratio 0.00), was not detected (n.d.) as a high-confidence G44D interactor in our AP-MS data. (B) Core Mediator subunits are retained in both studies (median G44D/WT 0.92 in Turunen et al., 0.87 in the present study). The selective reduction of the CDK module with retention of core Mediator reproduces the central finding of Turunen et al. (2014).
__Reviewer comment: __It is unclear whether the system is appropriate for ER-dependent transcription. Endogenous ERα/ERβ levels are not given, and estrogen signaling is ligand-dependent, yet no estrogen treatment appears to have been included. Without receptor expression and ligand stimulation, the estrogen-reporter results are hard to interpret.
__Response: __This is a valid concern. The estrogen reporter is a defined dual-luciferase construct in which tandem estrogen response elements (ERE) drive firefly luciferase, and it does not include a co-expressed estrogen receptor. Because HEK293 cells express very low to negligible endogenous ERα and the assays were performed without ER co-expression or 17β-estradiol stimulation, the readout reflects ERE-responsive promoter activity in a near-ER-null context by design, rather than bona fide ligand-dependent ER signaling. We now (i) report the endogenous receptor status directly: neither ESR1 nor ESR2 is detected among the 8,708 protein groups quantified in the parental Flp-In T-REx 293 line by Orbitrap Astral DIA (Supplementary Data 5), confirming an effectively ER-null background; and (ii) explicitly qualify the estrogen-reporter interpretation in both the Results and the Discussion. Repeating the assay with ERα co-expression and 17β-estradiol stimulation would be required for a ligand-dependent readout, and we agree this is the appropriate next experiment, but it falls outside the scope of the present revision and we have therefore limited our claims accordingly. The phosphorylation-dependent PELP1 association (PL-specific to Mut1) is presented as motivating, not demonstrating, an ER-signaling link.
Reviewer comment: Page 15, the authors described "Among individual kinases, HIPK2 was the top candidate for PAGE4-T51/T85, both sites fitting their known serine/threonine consensus motifs.". HIPK2 did not appear in the list of kinases in Figure 3E. Please clarify the data source and the criteria for the ranking.
__Response: __Reconciled as described for Reviewer 2-C: the kinase-family enrichment heatmap has been replaced by a site-level kinase-prediction panel for S9, T51 and T85 (lower panel of Figure 4A), figure and text name the same kinases, and the scoring method, ranking criterion and data source are stated in the caption and Methods.
Reviewer comment: Supplementary Figure 4, significance marks were missing.
__Response: __Significance annotations have been added to all panels of the luciferase figure, which is Figure S5 in the revised manuscript (Figure S4 is the pathway-enrichment heatmap). Each panel now shows the individual replicate values with the mean ± SD, and the legend states the test, the number of replicates, the comparator (all comparisons against PAGE4 WT) and the meaning of each asterisk level.
__Reviewer comment: __The biological significance of the enriched motifs is somewhat overinterpreted, since motif enrichment alone does not establish functional involvement. The presentation could be improved with a figure of top motifs, enrichment statistics and/or logos; e.g. “PAGE4 WT maintained AP-1-driven transcription, whereas phosphorylation induced an expanded motif repertoire” is hard to follow.
__Response: __We agree, and we have reworked the motif analysis to lead with effect size rather than statistical significance, correcting several over-statements in the original text.
__The p-values overstate the enrichment. __Across the five conditions, 129–172 of 440 known motifs pass q ≤ 0.05, but ~80% of these have fold-enrichment __The individually named transcription factors were over-called. __Most factors named in the original text reach statistical significance yet have fold-enrichment close to 1.0 (e.g. AP-1 1.03–1.11, E2F3 1.03, ETS1 1.07, FOXA1 1.03–1.09), i.e. no meaningful enrichment. We have removed these factor-specific claims and the “AP-1-driven transcription” / “expanded motif repertoire” framing.
CTCF is the one robust result. CTCF is the only motif with fold-enrichment ≥ 1.3 in all five conditions (range 1.38–2.29, peaking at 2.29 in PAGE4 S9D); its paralog BORIS is second-strongest (1.18–1.58). No other known motif exceeds 1.5-fold enrichment in more than one condition. We therefore name CTCF as the principal candidate regulatory motif, reported with effect sizes, and provide a summary figure of the top enriched motifs by fold-enrichment (Rebuttal Figure 3, below).
Rebuttal Figure 3. Top enriched known motifs (HOMER) by fold-enrichment across the five conditions. Colour intensity = fold-enrichment (% target / % background); grey dash = motif not detected in that condition; n.s. = not significant (HOMER Benjamini q > 0.05). CTCF is the only motif with fold-enrichment >= 1.3 in all five conditions.
Venn filter and counts (corrected). Consistent with the effect-size-first approach described above, the Venn diagrams in Figure 6E now show robustly enriched known motifs (HOMER Benjamini q ≤ 0.05 AND fold-enrichment ≥ 1.2); we recomputed the overlaps directly from the five HOMER outputs using this filter (union denominator). Corrected values — MED12: WT = 10, G44D = 18, union = 21, G44D-unique = 11 (52.4%), shared = 7 (33.3%); PAGE4: WT = 19, S9D = 12, S9A = 11, union = 25, shared by all three = 3 (12.0%), S9D-unique = 4 (16.0%), S9A-unique = 1 (4.0%). This supersedes our earlier q ≤ 0.05-only statement, which counted many weakly enriched motifs (fold-enrichment __Reviewer comment: __The Figure 6D data do not sufficiently support the statement that MED12 WT enriched for chromatin remodeling/transcriptional regulation, G44D for stress-associated functions, and phospho-PAGE4 for nucleosomal DNA binding and ribosomal interaction terms. Please clarify.
__Response: __We agree — the original panel over-interpreted small per-construct differences by assigning a distinct function to each construct. We have replaced the panel (now Figure 6G) with a corrected analysis that no longer makes per-construct functional assignments.
The new Figure 6G is a term-by-condition dot plot of Gene Ontology Molecular Function enrichment of the peak-associated genes (clusterProfiler, redundant terms reduced by rrvgo at similarity 0.9, Benjamini–Hochberg-adjusted p-values). It shows that all five constructs enrich for the same core terms: the most strongly enriched term in every condition is structural constituent of ribosome, and the other enriched terms (cadherin binding, rRNA binding, ribonucleoprotein-complex binding, ubiquitin-ligase-related terms) are likewise shared across conditions rather than unique to any one. No construct-specific functional classes were detected, consistent with RNAP II promoter occupancy at highly transcribed genes. The Results text and Figure 6G legend have been rewritten to describe this shared-core finding, and the earlier “differentially enriched across constructs” wording and the stale “Figure 6D” reference have been removed.
__Reviewer comment: __The legend says AP (light) vs PL (dark), but the figure shows AP green and PL yellow. And on the 73% statement, if the labels are correct, AP HCIs are 25 for Mut1 and 149 for WT, i.e. only 17% retained.
__Response: __The legend has been corrected to the actual colors used in the figure, and the retention value has been recomputed and corrected as described for Reviewer 2-D.
__Reviewer comment: __The data supporting “interactions with basal transcription factors and cell-cycle regulators were preferentially lost, while Mediator subunits and RNA-processing factors were selectively retained” are not readily apparent. Indicate the relevant figure/table.
__Response: __We now cite the specific evidence, the focused interaction matrix (revised Figure 4E) and Supplementary Dataset 3, and have added a categorized lost-versus-retained summary so the statement is directly traceable to the data.
__Reviewer comment: __Page 19, it is unclear how the statement that "In the MED12 p.G44D dataset, the bait (MED12-G44D) itself is strongly reduced compared with MED12 WT, consistent with decreased recovery of the mutant complex." is supported by the data presented in Figure 5A. Please clarify.
__Response: __Addressed as for Reviewer 2-E: the data show the MED12-G44D bait is recovered at levels equal to or higher than WT, so the original “bait reduced” statement was corrected; the reduced number of G44D interactors is retained as a genuine, bait-independent effect.
__Reviewer comment: __Page 19, reference is needed for "PELP1 is a well-established scaffolding protein that couples estrogen receptor signaling to chromatin remodeling and has been implicated in hormone-dependent tumor progression in breast and ovarian cancers."
__Response: __A reference for PELP1 as an estrogen-receptor coactivator/scaffold implicated in hormone-dependent tumors has been added.
__Reviewer comment: __For luciferase reporter assay: In the materials and methods section, the authors wrote "Data are presented as mean {plus minus} standard deviation (SD)", but in the figure legend, the authors wrote "Data represent mean {plus minus} SEM". Please clarify this discrepancy. In addition, please provide information regarding the promoter used in the luciferase assay plasmids, and replace "Firefly luciferase reporter plasmid containing the promoter of interest" with the specific promoter name.
__Response: __We have harmonized the error-bar reporting: the Methods and the figure legend now both state mean ± SD, with the same number of replicates given in each. We have also replaced “promoter of interest” with the specific response element for each pathway (cell cycle/pRb-E2F → E2F; p53/DNA-damage → p53; MAPK/ERK → serum response element/SRE; MAPK/JNK → AP-1; Wnt → TCF/LEF; estrogen → estrogen response element/ERE).
__Reviewer comment: __“Buyukcelebi K et al, 2023” is cited as both 2023a and 2023b; the same issue affects “Lin et al, 2018” and “Turunen et al, 2014.”
__Response: __We have consolidated the duplicated entries and made the year-suffix usage consistent for these and all other multiply-cited references. We also carried out a complete audit of the bibliography against the in-text citations: every citation now resolves to exactly one entry, no entry is duplicated, and no entry is left uncited. Several references that were missing from the list have been added, and two citations that could not be traced to a specific source were removed from a sentence that is adequately supported by the remaining reference. Prompted by this comment we also checked each citation against the statement it supports rather than only against the bibliography, and corrected several placements: two extracellular-matrix references have been moved onto the matrix sentence they actually support, a reference on the co-occurrence of MED12 and HMGA2 alterations has been moved to the sentence reporting concurrent MED12 mutations in the COL4A5-COL4A6 tumors, a PAGE4-specific reference has been added alongside the general androgen-receptor citation, and one claim in the Discussion has been narrowed to the extracellular-matrix and cytoskeletal remodeling that the cited work actually addresses.
__Reviewer comment: __“PAGE4 emerged … at protein, transcript, and IHC levels …”, IHC characterizes protein expression; did the authors mean phosphorylation levels?
__Response: __Corrected. IHC measured PAGE4 protein abundance and (cytoplasmic-predominant) localization; the sentence has been rewritten so the three independent layers are stated accurately (protein by MS, transcript by RNA-seq, and protein by IHC), without implying IHC measured phosphorylation.
__Reviewer comment: __“COL4A5-COL4A6 subclass showed upregulation of several extracellular matrix markers (WNT4, COL3A1, COL4A1, MMP2)”, it is unclear whether WNT4 should be an ECM marker.
__Response: __We agree and have reclassified WNT4 as a hormone/Wnt-signaling factor rather than an ECM component, and corrected the sentence so only bona fide ECM markers are grouped together.
-
Note: This preprint has been reviewed by subject experts for Review Commons. Content has not been altered except for formatting.
Learn more at Review Commons
Referee #3
Evidence, reproducibility and clarity
Uterine leiomyomas (UL) occur in more than 70% reproductive age women and mutations in the RNA Polymerase II (RNAP II) transcriptional Mediator subunit MED12 account for ~70% of these tumors. Despite this genetic insight, the molecular mechanisms linking MED12 mutations to UL pathogenesis remain incompletely understood, hindering the development of effective therapeutics to the disease. Through unbiased multi-omics analyses, including proteomics, phosphoproteomics, and interrogation of previously published RNA-seq datasets, followed by IHC validation, the authors identified a specific upregulation of PAGE4 protein expression, …
Note: This preprint has been reviewed by subject experts for Review Commons. Content has not been altered except for formatting.
Learn more at Review Commons
Referee #3
Evidence, reproducibility and clarity
Uterine leiomyomas (UL) occur in more than 70% reproductive age women and mutations in the RNA Polymerase II (RNAP II) transcriptional Mediator subunit MED12 account for ~70% of these tumors. Despite this genetic insight, the molecular mechanisms linking MED12 mutations to UL pathogenesis remain incompletely understood, hindering the development of effective therapeutics to the disease. Through unbiased multi-omics analyses, including proteomics, phosphoproteomics, and interrogation of previously published RNA-seq datasets, followed by IHC validation, the authors identified a specific upregulation of PAGE4 protein expression, phosphorylation at residues T51 and T85, and transcript abundance in MED12 mutated UL. Using Flp-In{trade mark, serif} 293 T-Rex cell model system, the authors performed affinity purification and proximity labeling assays to characterize the interactomes of different phosphorylation variants of PAGE4 and G44D mutant MED12. They also conducted ChIP-seq analyses to assess whether these mutations alter RNAP II occupancy on chromatin and performed luciferase reporter assays to determine their effects on the activities of major signaling pathways, including pRb, p53, MAPK/ERK, MAPK/JNK, Wnt, and estrogen. They concluded that PAGE4 phosphorylation acts as a molecular switch, dramatically remodeling its protein interaction landscape to favor associations with the Mediator complex and core transcriptional machinery. Their findings establish phospho-PAGE4 as a central mechanistic node in MED12-mutant leiomyoma pathogenesis, expand the biological role of PAGE4 beyond prostate cancer to encompass female reproductive tract neoplasia, and identify a promising biomarker and potential therapeutic target for the most common molecular subtype of UL. The presented findings are supported by the experimental data. The manuscript is well written and provides sufficiently detailed methods to ensure experimental reproducibility. Overall, the study is interesting and important and warrants publication. However, I do have several specific concerns that need to be resolved.
Major comments:
- For assays of interactome, luciferase reporter, and RNAP II occupancy on chromatin of PAGE4 phosphorylation variants and MED12 G44D mutants, one critical limitation is that the studies exclusively relied on tagged, exogenously expressed proteins without evaluation of the levels of endogenous proteins. The endogenous protein could substantially affect the observed interaction networks, particularly if PAGE4 and MED12 expressed at appreciable levels in the Flp-In{trade mark, serif} 293 T-Rex cells. I recommend that the authors measure endogenous protein expression levels in these cells and discuss how endogenous protein levels may influence interpretation of the data. Did different phosphorylation variants of PAGE4 affect their nuclear localization? It would also be interesting to compare the T51 and T85 PAGE4 phosphorylation status between wild-type MED12 and G44D MED12 cells.
- Turunen et al (2014, PMID: 24746821) reported mutant MED12 interactomes using the similar strategy. It would be valuable for the authors to compare their findings with this published dataset. Such a comparison could help assess the reproducibility and robustness of the identified protein-protein interactions and provide additional confidence of the current results.
- The authors used luciferase reporter assays to assess the effects of PAGE4 phosphorylation variants and G44D MED12 on multiple signaling pathways, including estrogen signaling. However, it is unclear whether the experimental system is appropriate for evaluating estrogen receptor (ER)-dependent transcriptional activity. Specifically, the manuscript did not provide information regarding the endogenous expression levels of ERα and/or ERβ in the cells. In addition, estrogen signaling is typically ligand-dependent, yet no estrogen treatment appears to have been included in the experimental design. Without demonstrating receptor expression and appropriate ligand stimulation, it is difficult to interpret the negative or positive effects of mutant proteins on estrogen-responsive reporter activity. Inclusion of information of estrogen receptor expression levels and ligand stimulation experiments would substantially strengthen the conclusions regarding the role of these variants in estrogen signaling.
Minor comments:
- Page 15, the authors described "Among individual kinases, HIPK2 was the top candidate for PAGE4-T51/T85, both sites fitting their known serine/threonine consensus motifs.". HIPK2 did not appear in the list of kinases in Figure 3E. Please clarify the data source and the criteria for the ranking.
- Supplementary Figure 4, significance marks were missing.
- For the motif enrichment section, while the findings are potentially interesting, the manuscript devotes substantial discussion to the presumed roles of these motifs in UL pathogenesis based solely on motif enrichment analysis. In my opinion, the biological significance of these motifs is somewhat overinterpreted, as motif enrichment alone does not establish functional involvement of the corresponding transcription factors in tumor development or progression. In addition, the presentation of the motif analysis could be improved. I recommend including a figure summarizing the top enriched motifs, along with enrichment statistics and/or motif logos. A visual representation would make the results more accessible and help readers better appreciate the key findings without relying exclusively on lengthy textual descriptions. For example, it is difficult to understand the statement that "PAGE4 WT maintained AP-1-driven transcription, whereas phosphorylation induced an expanded motif repertoire."
- The data presented in Figure 6D did not sufficiently support the statement that "MED12 WT enriched for chromatin remodeling and transcriptional regulatory functions, G44D enriched for stress-associated functions, and phosphorylated PAGE4 mutants enriched for nucleosomal DNA binding and ribosomal interaction terms". Please clarify.
- Figure 4B: In Figure legend, the authors wrote "Stacked bars distinguish interactions captured by AP (light) versus PL (dark).". But AP was labeled as green and PL was yellow on Figure 4B. Page 76, the authors described "Among the AP-derived HCIs, the phosphomimic triple mutant (Mut1: S9D/T51E/T85E) retained only 73% of the interactors captured by PAGE4 WT, representing a net loss of approximately 27% of stable associations". If labels on Figure 4B were correct, this reviewer's understanding is that "In AP, HCIs for MUT1 is 25, and HCIs for WT is 149. Therefore, only 17% (25/149) stable interaction was retained". Please clarify.
- Page 16, The data supporting the statement that "This reduction was not uniform: while interactions with basal transcription factors and cell-cycle regulators were preferentially lost, associations with Mediator subunits and RNA-processing factors were selectively retained." are not readily apparent. Please indicate the relevant figure, table, or dataset.
- Page 19, it is unclear how the statement that "In the MED12 p.G44D dataset, the bait (MED12-G44D) itself is strongly reduced compared with MED12 WT, consistent with decreased recovery of the mutant complex." is supported by the data presented in Figure 5A. Please clarify.
- Page 19, reference is needed for "PELP1 is a well-established scaffolding protein that couples estrogen receptor signaling to chromatin remodeling and has been implicated in hormone-dependent tumor progression in breast and ovarian cancers."
- For luciferase reporter assay: In the materials and methods section, the authors wrote "Data are presented as mean {plus minus} standard deviation (SD)", but in the figure legend, the authors wrote "Data represent mean {plus minus} SEM". Please clarify this discrepancy. In addition, please provide information regarding the promoter used in the luciferase assay plasmids, and replace "Firefly luciferase reporter plasmid containing the promoter of interest" with the specific promoter name.
- The same ref "Buyukcelebi K et al, 2023" was cited differently as 2023a and 2023b. The same mistakes for refs. "Lin et al, 2018" and "Turunen et al, 2014".
- Page 25, the authors wrote "PAGE4 emerged as a consistently induced marker in MED12-mutant ULs at protein, transcript, and IHC levels, with cytoplasmic-predominant staining that cleanly distinguishes UL from matched myometrium". The IHC level characterized protein expression. Did the authors mean "phosphorylation levels" rather than "IHC levels"?
- Page 13, the authors wrote "Besides PAGE4, COL4A5-COL4A6 subclass showed upregulation of several extracellular matrix markers (WNT4, COL3A1, COL4A1, and MMP2)". It is unclear whether WNT4 should be categorized as an ECM marker. The authors should provide a rationale for the classification or consider using a more appropriate designation.
Significance
The authors convincingly demonstrated that PAGE4 mRNA and protein levels and the levels phosphorylation at T51 and T85 are specifically elevated in MED12 mutant UL. Although PAGE4 is traditionally regarded as a male-specific marker of prostate cancer and other urogenital diseases, the authors found it to be markedly upregulated and hyperphosphorylated in these female-specific tumors. This unexpected observation suggests that PAGE4 may have a broader, sex-independent function in cellular signaling pathways and UL progression. One limitation of this study is the absence of functional investigations examining the effect of PAGE4 on cell proliferation, apoptosis, or other UL-associated phenotypes, which would be important for establishing its biological relevance. In contrast, much of the manuscript focuses on interactome (affinity purification and proximity labeling approaches) or transcriptional activity (ChIP-seq analysis of RNAP II chromatin occupancy and luciferase-based transcriptional reporter assays) analyses of mutation variants in a non-disease-relevant cell model, making it difficult to assess the significance of these findings for UL pathogenesis. Consequently, the mechanistic conclusions remain preliminary without functional validation in biologically relevant UL models. Nevertheless, this study will be of significant interest to UL researcher community, stimulating further investigation into the role of PAGE4 in UL cell survival and growth, and supporting translational efforts to evaluate its utility as a biomarker or and potential therapeutic target.
-
Note: This preprint has been reviewed by subject experts for Review Commons. Content has not been altered except for formatting.
Learn more at Review Commons
Referee #2
Evidence, reproducibility and clarity
The manuscript by Bong et al. describes a comprehensive proteogenomic analysis of uterine leiomyomas tumors (ULs). The goal was to identify protein level effectors that drive MED12 mutant UL tumors. They employed discovery DDA global and phosphoproteomic analyses of 9 MED12-mutated UL and matched normal tissues. They identified that PAGE4 was upregulated at both the protein and phosphosite level. They further characterized MED12 ULs against other subtypes of UL (HMG2A, FH, COL4A) and confirmed PAGE4 expression and phosphorylation is confined to the MED12-mutant subtype. They constructed isogenic cell lines (HEK293) of PAGE4 and …
Note: This preprint has been reviewed by subject experts for Review Commons. Content has not been altered except for formatting.
Learn more at Review Commons
Referee #2
Evidence, reproducibility and clarity
The manuscript by Bong et al. describes a comprehensive proteogenomic analysis of uterine leiomyomas tumors (ULs). The goal was to identify protein level effectors that drive MED12 mutant UL tumors. They employed discovery DDA global and phosphoproteomic analyses of 9 MED12-mutated UL and matched normal tissues. They identified that PAGE4 was upregulated at both the protein and phosphosite level. They further characterized MED12 ULs against other subtypes of UL (HMG2A, FH, COL4A) and confirmed PAGE4 expression and phosphorylation is confined to the MED12-mutant subtype. They constructed isogenic cell lines (HEK293) of PAGE4 and MED12 variants to characterize the interactomes of PAGE4 and MED12. These results showed PAGE4 is involved in the mediator complex functions of MED12.
Luciferase reporter assays were used to characterize how PAGE4 and MED12 mutants affect the transcription of key genes.
They performed ChIP-seq experiments for RNA Pol2 to characterize how PAGE4 and MED12 mutants affect Pol2 transcription.
Overall, Bong et al. convincingly identified a novel role of PAGE4 in UL and mediator transcriptional processes. Thus, this work will represent an important novel contribution to understanding the pathobiology of UL.
There are several critiques for this manuscript. Major critiques:
A) The validation cohort is not technically a validation cohort as it uses the same 9 MED12 mutant UL samples. Thus, it cannot validate PAGE4 upregulation and phosphorylation in an independent cohort. A second MED12-mutant UL cohort should be used.
B) The ChIP-seq data are over interpreted. No meaningful changes in peak width or height is readily observable. To be believable statistical tests would be required.
C) HIP2K was proposed as a candidate kinase that phosphorylates PAGE4. This is based on data presented in Fig. 3E. However, there is no HIP2K on this heatmap. Instead CLK2 is highlighted in red font and has an increased score, indicating it could be responsible kinase. HIP2K and CLK2 are different unrelated kinases. Please rectify.
D) For Fig. 4B the text states MUT1 retained 73% of prey proteins contained in the PAGE4 WT interactome. However, in Fig. 4B, MUT1 had 25 preys versus 149 preys for WT; this is 25/149 = 17% retained for a decrease of 83%. Please rectify.
E) The text states the MED12-G44D bait protein is strongly reduced compared to MED12 WT, referring to Fig. 5A. Fig. 5A right panel does not support this conclusion. Please rectify.
F) In Fig. S3 the text states MUT1 has upregulated processes for base-excision repair and SUMOylation. However, the Fig. S3 heatmap does not support this for Mut1. Instead the data support MUT6 not MUT1. Please rectify.
G) Figure 4E currently shows the full bait-prey graph and becomes a "hairball" where edge density prevents the reader from seeing the differences claimed in the result section. This figure is uninterpretable. The bait protein names are in miniscule point font and are illegible. There are too many elements to see the difference between AP, PL, or Both or the name of the prey proteins. The edge lines widths and colors are uninterpretable. Suggest simplifying the visualization or split it into focused panels so the figure evidence matches the claims. Specifically, recommend aggregating preys by annotated complex (or functional group) and showing a single summarized edge between each bait and complex (with edge weight or thickness reflecting the number or strength of prey links), and then displaying only the individual preys that are differentially associated between baits to support the key statements. H) Fig. 5A presents the known phosphorylation sites for PAGE4. However, the figure shows the data for 'GAGE', which is a different (unused) name for PAGE4. Please rectify. Importantly, this figure is directly copied from PhosphoSitePlus®.org without apparent permission. Permission should be obtained.
I) The discussion and conclusions suggest making phosphomimetic changes at key sites in PAGE4 and MED12 are conclusive for phosphorylation being involved. This is not formally true, and should at least be mentioned as a hypothesis.
J) In general, the methods describing the mass spectrometry experiments need to be expanded and clarified. The current description is insufficient for reproducing the experiments or fully understanding how the different datasets were generated and analyzed. DIA MS methods are missing entirely, and given that the manuscript uses both DDA and DIA acquisition modes across multiple MS experiments (total proteomics, phosphoproteomics, AP, and PL), the methods should be reorganized to clearly explain how acquisition and subsequent analyses were performed for each experiment type, specifying what was done in DDA and what in DIA modes. In the downstream analysis, how were the total and phosphoproteomics normalized and integrated? It is not clear how the PL data were analyzed or whether they were treated the same way as AP data; please explicitly state whether PL and AP data were processed identically or with different pipelines and describe any differences in preprocessing, normalization, or statistical treatment. Was SAINT used on spectral counts or intensities? Since the baits are nuclear proteins, was anything with Nuclear Localization Signals used as controls?
K) Overall the manuscript has many issues with data presentation. To illustrate just a few:
- Fig. 3B and 3D are described as volcano plots. They are not volcano plots.
- Fig. 3B legend mentions a lower panel. There is no lower panel for Fig. 3B.
- Fig. 3A, please use different shape markers for the UL subtypes and suggest connecting with edges to thus support their text conclusions of separate groupings.
- For Fig. 3E the colors for the families and groups are not distinguishable (and the font is too small). Please rectify.
- Please label the units and scale for Fig. 6A.
- Please mention the MAC3 tag uses Ultra-ID to thus justify the 10 min labeling time.
- In the relevant protein and phosphoproteomic datasets and figures please mention the number of proteins and phosphosites obtained to thus allow an interpretation of the data quality.
- Table 4 is in a format that is uninterpretable.
- The text and figures interchangeably use the mutant number (e.g. MUT1) or mutant composition (e.g. S9D_T51E_T85E - for MUT1), this makes reading and interpreting the manuscript very challenging. Suggest using one format.
Significance
Overall, Bong et al. convincingly identified a novel role of PAGE4 in UL and mediator transcriptional processes. Thus, this work will represent an important novel contribution to understanding the pathobiology of UL.
This is a very comprehensive dataset. However, there are multiple issues on data interpretation and presentation.
The results would be of interest to cancer biologists, primarily at the basic science level. Clinical translation would require additional research.
The expertise of the reviewer is in cancer proteomics (interactomes, DIA and DDA mass spectrometry, proteomic data analysis), and general cancer biology across most common cancer types.
-
Note: This preprint has been reviewed by subject experts for Review Commons. Content has not been altered except for formatting.
Learn more at Review Commons
Referee #1
Evidence, reproducibility and clarity
Bong Yih Tyng and Xiaonan Liu employed a multi-omics proteomics approach (full proteome profiling, phosphoproteomics), mining of external datasets (RNAseq, ProteomeDB, Protein Atlas and an experimental validation strategy interactome mapping, proximity interactome mapping) to establish PAGE4 phosphorylation status as driver in MED12-mutant uterine leiomyomas. They used state-of-the-art DIA proteomics approaches to investigate in addition to MED12-variant tumors the PAGE4 expression in other tumors (COL4A5-COL4A6, HMGA2, and FH subclasses) substantiating their initial findings in MED12-variant tumors. The analysis strategies employed …
Note: This preprint has been reviewed by subject experts for Review Commons. Content has not been altered except for formatting.
Learn more at Review Commons
Referee #1
Evidence, reproducibility and clarity
Bong Yih Tyng and Xiaonan Liu employed a multi-omics proteomics approach (full proteome profiling, phosphoproteomics), mining of external datasets (RNAseq, ProteomeDB, Protein Atlas and an experimental validation strategy interactome mapping, proximity interactome mapping) to establish PAGE4 phosphorylation status as driver in MED12-mutant uterine leiomyomas. They used state-of-the-art DIA proteomics approaches to investigate in addition to MED12-variant tumors the PAGE4 expression in other tumors (COL4A5-COL4A6, HMGA2, and FH subclasses) substantiating their initial findings in MED12-variant tumors. The analysis strategies employed are suitable and mostly well-described (see major and minor comments regarding some points which might can be improved). The interaction proteomics and proximity labelling interactomics capture of PAGE4 WT and phosphomimics PAGE4 mutants (Mut1 triple phospho, Mut6 T51E mutant) and the MED12-G44D variant reveal extensive network rewiring towards RNA processing, ribosome biogenesis, translation and protein folding, with less interactions linking to its core transcriptional function. Together the study offers insights into the functional consequences of Med12-mutants uterine leiomyomas and establishes PAGE4 phosphorylation as a key factor for the network rewiring leading to adaptions in cell cycle progression, proliferation, and cellular stress responses.
Major comments:
Figure 1B: The first PCA component is mostly dominated by a single sample far to the right dominating the first principal component (33.9%). Was there an outlier test (such as Mahalanobis Distance or Z-score on PC1) performed to ensure that subsequent statistical analysis is not affected (e.g. skew group means, inflation of variance) by this single measurement? If an outlier test was conducted it should be reported in the method section. Also for the phosphoproteomics no PCA or outlier test is reported to judge data and analysis strategy this would be necessary. The full proteome and phosphoproteomics data need more details regarding the analysis to ensure reproducibility. Figure 3A: Full proteome PCA: please perform an outlier test for the HMGA2 on the left side. Phosphosites PCA: please perform an outlier test for the FH_UL and COL4A5-COL4A6_UL data points which are visually separated from the other groups. Like this the PCA are difficult to interpret as the data points for example in the full proteome PCA get squeezed together on an axis (PC1 mainly is driven by one measurement). In the manuscript the outliers are discussed showing inter-tumor heterogeneity, but it could be as well origin from variability in tissue collection or MS-performance. Outliers might affect statistics and conclusions regarding the statement about variation within the group. This should be tested if discussed in the manuscript.
Minor comments:
Phosphoproteomic analysis identifies aberrant kinase networks and PAGE4 hyperphosphorylation (page 8). Why is SNRNP70_S226 mentioned as being one of the most significantly hyperphosphorylated sites in UL. According to table S1 it has a log2FC of 5.92 and a p-value of 0.02 and was classified as non-significant. The data in table S1 contradicts the statement in the manuscript. The other example MCM2_S139 is significantly dysregulated. Also, MARCKSL1 (log2FC of 2.14 and p-value of 0.016, assessed as "Non-significant") also for NDRG2. The volcano plot and figure caption indicate a p-value cutoff of 0.05 instead of 0.01 which is apparently shown in table S1. Please clarify.
Phosphoproteomic analysis identifies aberrant kinase networks and PAGE4 hyperphosphorylation (page 8). Add citations for Protein Atlas and ProteomicsDB.
PAGE4 expression is selectively enriched in MED12-mutant UL (page 14): "PAGE4 expression indeed was significantly elevated in MED12-mutant ULs ... ". The figure caption, methods and text do not provide any information of the statistical test used regarding the statement of significant upregulation. Please provide the data or change wording as a statistical significance is implied.
Method section "Liquid chromatography-mass spectrometry (LC-MS)" report the UniProtKB database download date to ensure reproducibility. Does the database contain isoforms and TrEMBL sequences or uses a filtered version (e.g.: status reviewed, canonical?) (report it like on p.36 for the AP-MS/BioID samples). Further add a statement if missing values were imputed on which level (peptide, protein), which normalization strategy was employed, and was there filtering of incomplete values performed. All these analysis steps are common practice for DDA measurements of clinical samples and may affect subsequent statistical finding and thus should be reported in the method section if they were or were not performed (for the PPI data this was reported on p.37, but for the full proteome and phosphoproteome the information provided is not sufficient to judge the validity of data processing). Method section "GO enrichment and statistical analysis" (p.37): Please add citations for the used tools SAINT, CRAPome and Enrichr.
Figure 1c: The scale of the (Gene) count does not correspond with the circle sizes of the plot. It is thus not possible to judge the number of genes per category. This should be sized equally as else it is not possible for a reader to judge the number of genes per enriched terms. Figure 1G: Expression levels for ProteomicsDB. As the log2 of the expression level is used, it is not clear if the white squares correspond to missing values in the ProteomicsDB, or if the value was very low. Maybe indicating absent values in a grey color scheme makes this more explicit. Figure 1C: the caption indicates a "gray box" but it seems that PAGE4 was highlighted as a red dot, and a white box was added for indicating the gene name.
Figure 3D: in the manuscript (p.15) the plot is described as volcano plot. This looks like a scatterplot with horizontal scattering and indication of the significant threshold than a volcano plot (in principle it shows the data of the volcano plot, but one expects the adjusted p-value as axis in a classic volcano).
Figure 5E: The PPI network is difficult to read as there are too many edges and nodes. Optional: It could make sense to provide a more refined analysis (collapsing the categories/ proteins) as like this it is impossible to catch differences or provide a subnetwork and provide the full network in the supplementary.
Significance
The identification of PAGE4 in MED12 mutant ULs, which is typically expressed (see Figure S1) in male as an upregulated and hyperphosphorylated protein in leiomyomas is a novel finding. The sex-independent role was clearly supported by various data layers (proteome and phosphoproteome screen, IHC and secondary DIA validation dataset). The strength of the study is that they have multiple layers of protein evidence for this specific finding. In addition, they performed additional functional genomics and interactomics experiments to elucidate mechanistic understanding and consequences of PAGE4 hyperphosphorylation and identified a phosphorylation switch and were able to link HIPK2 as likely kinase. These findings are appropriately highlighted and discussed in the study. The study does provide a deeper understanding of uterine leiomyomas tumors.
The study links PAGE4 expression and phosphorylation to MED12-mutant ULs. It provides levels of evidence for this on the mRNA, protein and phosphoproteome. It establishes that PAGE4 is MED12-mutant specific using a second cohort with different cancer markers. The techniques (proteomics, phosphoproteomics, interactomes) are conducted on an expert level. The data analysis is conducted at sufficient level, only outlier tests for samples which do not follow the trends are lacking as a major criticism. The reporting of the results is appropriate. A strength is that the phosphoproteome data was of high quality to provide side-specific information and employ normalization using the unphosphorylated abundance. I recommend providing outlier tests and providing more detail on the method section for the initial DDA measurements.
Audience:
The study is of broad interest, due to the prevalence of the uterine leiomyomas. Further, the combined proteomics/ interactomics approach combined with functional validation will be appealing for several fields including systems biology, clinical/translational, and molecular biology orientated researchers. The research might serve as starting point for larger clinical proteomics or biomarker studies. Further studies might focus on strengthening the validation of HIPK2 or investigating the link to PELP1 upon phosphorylation of PAGE4.
My expertise:
My field of expertise are molecular systems biology, full proteome profiling, and MS-based interaction proteomics analysis.
-
