TEM-LLM

TEM-LLM is a domain-specific multimodal large language model for Transmission Electron Microscopy (TEM) image analysis, fine-tuned from LLaVA-v1.5-7B using a difficulty-aware curriculum learning strategy.

This model is associated with the following work (currently under review):

AI-Assisted Materials Characterization: A Curriculum-Guided Multimodal Framework for Transmission Electron Microscopy


Model Details

Base Model LLaVA-v1.5-7B (Vicuna-v1.5-7B)
Training Strategy Difficulty-Aware Curriculum Learning (4 stages)
Training Data ~216K QA pairs across 40K TEM images
Domain Transmission Electron Microscopy (TEM)
Modalities CTEM, HR-TEM, STEM, Diffraction
Fine-tuning Method LoRA (Low-Rank Adaptation)
License LLaMA 2

Intended Use

TEM-LLM is designed for:

  • Image Classification: Identifying TEM imaging modalities (CTEM, HRTEM, STEM, diffraction)
  • Microstructural Feature Detection: Localizing and counting structural elements such as nanoparticles and grain boundaries
  • Feature Analysis & Interpretation: Providing scientific explanations of observed microstructural phenomena
  • Comprehensive Image Description: Integrating multiple aspects of visual understanding into structured descriptions

Important: Inference Requirements

TEM-LLM is built on LLaVA and cannot be loaded directly via transformers. Inference requires the LLaVA repository.

Step 1 โ€” Clone LLaVA:

git clone https://github.com/haotian-liu/LLaVA.git
cd LLaVA
pip install -e .

Step 2 โ€” Download weights:

from huggingface_hub import snapshot_download
snapshot_download(
    repo_id="LabSmart/TEM-LLM", 
    local_dir="./TEM-LLM"
)

Step 3 โ€” Run inference:

python -m llava.serve.cli \
    --model-path "./TEM-LLM" \
    --image-file "path/to/your/tem_image.jpg" \
    --load-4bit

For the complete pipeline including data crawling, preprocessing, knowledge distillation, training, and evaluation scripts, please refer to the GitHub repository.


Limitations

  • TEM-LLM is optimized for TEM image analysis and may produce unreliable outputs when applied to out-of-scope questions or out-of-domain images.
  • The model provides approximate qualitative descriptions rather than exact quantitative measurements.
Downloads last month
11
Safetensors
Model size
7B params
Tensor type
F16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for LabSmart/TEM-LLM

Finetuned
(33)
this model