A Conceptual Framework for Olfactory Perception in Large Language Models
Large language models (LLMs) have achieved multimodal perception across vision, audio, and text. Olfaction — the sense of smell — remains one of the major human sensory modalities without a corresponding digital input modality for LLMs.
This paper proposes the Molecular Olfaction Architecture (MOA), a conceptual framework in which a molecular detection layer identifies volatile organic compounds (VOCs) present in an environment and passes them as structured input to an LLM. The LLM then applies its learned chemical and semantic knowledge to produce a natural-language interpretation of the detected scent.
We hypothesize that LLMs already possess substantial implicit knowledge of chemistry and olfaction acquired during pretraining, and that MOA could provide the missing sensory bridge between physical molecular detection and semantic reasoning.
An informal proof-of-concept demonstrates the potential viability of the reasoning layer independently of physical sensing hardware. This paper describes the proposed architecture, its potential applications, limitations, and directions for future empirical validation.
The development of multimodal artificial intelligence has followed a relatively consistent pattern: connect a perception encoder to a language model, and the model gains the ability to reason about information originating from a new sensory modality.
Vision-language models such as GPT-4V and Gemini demonstrate this paradigm for visual information. Audio-language models extend similar capabilities to speech and environmental sound.
In these systems, the underlying language model does not necessarily need to be fundamentally redesigned to process a new modality. Instead, it can receive a structured representation of sensory information through an appropriate input interface.
Olfaction has not yet followed the same path.
Despite being one of the most chemically complex and information-rich human senses, smell does not currently have a standardized digital input modality integrated into general-purpose LLM systems.
Current LLMs can discuss smells and describe the expected odor of substances such as coffee, rain, gasoline, or flowers. However, they cannot directly perceive these odors from the physical environment.
The missing component is an olfactory encoder capable of converting molecular information into a representation that an LLM can process.
This paper proposes the Molecular Olfaction Architecture (MOA) as a conceptual solution.
MOA consists of a molecular detection layer that identifies volatile compounds present in the surrounding environment and an LLM reasoning layer that interprets those compounds semantically, producing a human-readable description of the corresponding olfactory profile.
The central hypothesis is that the semantic knowledge required for this interpretation may already exist within general-purpose LLMs.
If this is the case, the primary missing component is not necessarily a new model architecture, but rather a reliable sensory input channel.
Electronic nose (e-nose) technology has existed for decades. Traditional e-noses typically employ arrays of chemical sensors, including metal oxide semiconductor (MOS) sensors and conductive polymer sensors, to generate an electrical fingerprint associated with an odor.
These fingerprints are commonly processed using pattern-recognition and machine-learning techniques such as principal component analysis (PCA), support vector machines (SVMs), and other classification methods.
Although effective for specific applications, this approach is fundamentally limited by the scope of its learned classification space. An e-nose generally identifies odors according to previously defined classes or reference samples and produces a predefined classification output.
Such systems are not inherently designed to perform open-ended semantic reasoning about molecular compositions.
For example, an e-nose may identify a sample as "coffee," but it does not necessarily possess the ability to reason about why the sample smells like coffee, which compounds contribute to the perception, or what a novel combination of compounds might imply.
LLMs provide a potentially complementary capability.
Through pretraining on scientific literature, chemical databases, technical documentation, and general text, LLMs may acquire associations between chemical compounds and their known properties, including olfactory characteristics.
For example, an LLM may associate geosmin with the characteristic earthy odor commonly perceived after rain on dry soil, or associate 2-furfurylthiol with roasted coffee aroma.
This knowledge is normally latent and cannot be directly triggered by real-world molecular measurements because current LLM systems generally lack an olfactory sensory interface.
MOA proposes connecting these two domains:
MOA proposes a three-stage processing pipeline.
A chemical sensor array, electronic nose, gas chromatography system, or mass spectrometer samples the ambient air and estimates which volatile organic compounds (VOCs) are present.
This stage is analogous to an image encoder in a vision-language system. It converts a physical phenomenon into structured digital information that can subsequently be processed by an AI model.
The detected compounds are transformed into a standardized representation suitable for LLM processing. A simplified example:
2-Furfurylthiol: high Pyrazines: high Diacetyl: medium Guaiacol: low Acetic acid: low
A more advanced representation could include estimated concentrations, confidence scores, molecular identifiers, sensor reliability, and environmental metadata such as temperature and humidity. For example:
Compound: 2-Furfurylthiol Estimated concentration: high Detection confidence: 0.94 Compound: Pyrazines Estimated concentration: high Detection confidence: 0.88 Compound: Diacetyl Estimated concentration: medium Detection confidence: 0.81
The structured molecular representation is provided to an LLM as sensory input. The LLM applies its learned knowledge of chemistry, molecular associations, odor descriptors, and environmental context to infer a likely olfactory profile.
The key architectural hypothesis is that Stage 3 may not require a specialized olfactory language model or extensive fine-tuning. If general-purpose LLMs already contain sufficient chemical and olfactory knowledge, MOA primarily needs to provide a reliable sensory input channel capable of exposing that knowledge to real-world molecular data.
To evaluate the reasoning layer of MOA independently of physical sensing hardware, an informal proof-of-concept test was conducted. A list of volatile compounds associated with freshly brewed roasted coffee was manually composed and provided to a general-purpose LLM (Google Gemini).
"You are a test of a new architecture emerging for olfaction in LLMs. Identify this scent and I will tell you if you are correct: 2-Furfurylthiol, Geosmin, Diacetyl, Pyrazines, Acetic acid, Formic acid, Guaiacol, Furaneol."
The model identified the target scent as freshly brewed roasted coffee and provided a detailed interpretation of the potential contribution of the listed compounds to the overall olfactory profile.
The molecular input was manually constructed by a human who already knew the target scent. The compound list was not generated by an independent molecular sensor or spectrometry system.
Therefore, this experiment does not constitute empirical validation of the complete MOA pipeline. It should instead be interpreted as a preliminary demonstration that the semantic reasoning layer can accept a molecular representation and generate a plausible olfactory interpretation using an existing general-purpose LLM without architectural modification or task-specific fine-tuning.
A rigorous evaluation would require:
If implemented with sufficiently accurate and portable molecular sensing hardware, MOA could enable a class of applications that are currently difficult or impossible for conventional AI systems.
Instead of simply detecting a predefined chemical signature, the system could potentially reason about combinations of compounds and describe their meaning in natural language.
Human breath, skin emissions, and other biological samples contain volatile organic compounds that may correlate with physiological or pathological states. MOA could potentially assist researchers and clinicians by transforming detected volatile profiles into interpretable descriptions or hypotheses.
Rather than producing only a binary classification PASS / FAIL, a system could generate: "Detected profile is consistent with roasted coffee, with elevated sulfur-containing volatiles and pyrazines." This could support quality control, production monitoring, anomaly detection and product consistency analysis.
An olfactory interface could potentially provide individuals with anosmia or reduced olfactory perception with a digital representation of environmental smells — e.g. "Freshly cut grass.", "Strong citrus odor with a dominant lemon-like profile.", "Possible smoke detected."
Mass spectrometers can be expensive and difficult to miniaturize. Consumer-grade MOS arrays are accessible but detect broad responses rather than individual molecules. Trade-off: Cost ↔ Portability ↔ Molecular Specificity ↔ Detection Accuracy.
Aging, environmental conditions, contamination and changes in sensor characteristics require calibration and degradation compensation.
Real-world odors can consist of dozens or hundreds of VOCs. Unclear how accurately an LLM interprets complex mixtures with nonlinear interactions poorly represented in training data.
Presence alone does not determine perceptual importance. Human olfaction depends on concentration, thresholds, interactions, mixture effects and individual differences.
Confident but incorrect interpretations for unusual combinations, synthetic compounds, incomplete or conflicting profiles. Requires uncertainty estimation, chemical verification, RAG, structured validation and confidence propagation.
No complete experiment with real-time sensing, automated identification, structured generation and blind identification has been conducted. Feasibility remains an open empirical question.
Controlled experiments with defined scent corpus: identification accuracy, compound-to-scent reasoning, robustness to concentration, unseen combinations, false-positive rate, confidence calibration, agreement with human assessments.
RAG over chemical databases (structures, odor descriptors, thresholds, properties) to reduce hallucination and improve grounding.
Evaluate fine-tuning on olfactory literature, odor datasets, molecular-to-odor mappings and human assessments vs general-purpose LLMs.
The central proposal is that an AI system may not need to learn smell from scratch. Instead: a molecular sensing layer could convert real-world odors into structured chemical information, while an LLM could use its existing knowledge to interpret it.
The informal proof of concept suggests the reasoning layer is plausible. However, the major unresolved challenge is the sensory interface: developing affordable, portable, reliable and precise molecular detection hardware as an olfactory encoder.
In vision, cameras bridge world and model. In audio, microphones. MOA proposes molecular sensors could play an analogous role for olfaction — turning smell from something AI talks about into something AI can measure, interpret and reason about.