Verified Language Processing with Hybrid Explainability

Oliver Robert Fox
Giacomo Bergami
Graham Morgan

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

The volume and diversity of digital information have led to a growing reliance on Machine Learning (ML) techniques, such as Natural Language Processing (NLP), for interpreting and accessing appropriate data. While vector and graph embeddings represent data for similarity tasks, current state-of-the-art pipelines lack guaranteed explainability, failing to accurately determine similarity for given full texts. These considerations can also be applied to classifiers exploiting generative language models with logical prompts, which fail to correctly distinguish between logical implication, indifference, and inconsistency, despite being explicitly trained to recognise the first two classes. We present a novel pipeline designed for hybrid explainability to address this. Our methodology combines graphs and logic to produce First-Order Logic (FOL) representations, creating machine- and human-readable representations through Montague Grammar (MG). The preliminary results indicate the effectiveness of this approach in accurately capturing full text similarity. To the best of our knowledge, this is the first approach to differentiate between implication, inconsistency, and indifference for text classification tasks. To address the limitations of existing approaches, we use three self-contained datasets annotated for the former classification task to determine the suitability of these approaches in capturing sentence structure equivalence, logical connectives, and spatiotemporal reasoning. We also use these data to compare the proposed method with language models pre-trained for detecting sentence entailment. The results show that the proposed method outperforms state-of-the-art models, indicating that natural language understanding cannot be easily generalised by training over extensive document corpora. This work offers a step toward more transparent and reliable Information Retrieval (IR) from extensive textual data.

Version published to 10.3390/electronics14173490
Aug 31, 2025
Version published to 10.20944/preprints202504.0090.v2
May 16, 2025
Version published to 10.20944/preprints202504.0090.v1
Apr 2, 2025

Substitute-Space Embeddings for Label-Free Syntax: Unsupervised AI for POS Discovery

This article has 1 author:
1. Vipul Razdan
This article has no evaluationsLatest version Jan 8, 2026
Image and Video Question Answering with Large Language Models: A Comprehensive Review

This article has 3 authors:
1. Alexander Davis
2. Justin Parker
3. Julian Perry
This article has no evaluationsLatest version Dec 19, 2025
Variability in Low-Resource Machine Translation Evaluation: Authentic vs. LLM-Generated Training Corpora

This article has 3 authors:
1. Sofía García González¹
2. German Rigau Claramunt²
3. Jose Ramom Pichel Campos
This article has no evaluationsLatest version Jan 21, 2026

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

Substitute-Space Embeddings for Label-Free Syntax: Unsupervised AI for POS Discovery

Image and Video Question Answering with Large Language Models: A Comprehensive Review

Variability in Low-Resource Machine Translation Evaluation: Authentic vs. LLM-Generated Training Corpora