BidirLM Omni 2.5B, Core ML W8A16, Neural Engine and GPU, 8K context

Compiled Core ML programs for text, still-image, audio, and mixed-message embeddings, one per family (language, vision, audio), with a path for the Mac's Neural Engine and one for its GPU over each family's one set of INT8 weights:

  • Neural Engine: 97 staged functions; every operation was audited and actual Neural Engine execution was traced when the release was built.
  • GPU: one complete encoder per family (language, vision, audio front end and encoder), a whole item per prediction. FP16 matrix work with an FP32 residual stream; the faster mode for long inputs.

A message holds up to 8,192 tokens, chat template and media included, in either mode. Both modes embed into the same space: the GPU's vectors are gated against the Neural Engine's, so one index can hold vectors from either. There is no CPU mode.

Every message produces the same kind of vector: 2,048 dimensions, L2-normalized, masked mean pooling over all of the message's tokens (chat-template tokens included). There are no query or document prompts; queries and documents are encoded identically.

Source: BidirLM-Omni-2.5B-Embedding, revision 447a6e31be61b84443144afda21374339ce408e6. Projection weights use per-channel INT8 with FP16 activations; the token-embedding table stays FP16.

Releases

The same functions are published in these layouts, and every one embeds into the same space (the manifest's spaceID, ...:w8a16:ane-8k-v1), so one index can hold vectors from any of them:

Repository Modes Programs
1of2/bidirlm-omni-2.5b-coreml-w8 Neural Engine and GPU one program for every function (revision 57509034)
1of2/bidirlm-omni-2.5b-coreml-w8-bundle (this repository) Neural Engine and GPU one program per family
1of2/bidirlm-omni-2.5b-coreml-w8-ane Neural Engine one program per family
1of2/bidirlm-omni-2.5b-coreml-w8-gpu GPU one program per family

One program per family loads faster: Core ML reads a function's whole program each time it loads one. Revisions of 1of2/bidirlm-omni-2.5b-coreml-w8 up to c228769d56fad334c82d726c31f184122f04f5f9 held a withdrawn release (r4, 32,768 tokens) that embeds into the source model's space too but is not bit-identical: re-embed rather than mix its vectors into an index of these.

Requirements

  • An Apple-silicon Mac with macOS 15 or later (multifunction Core ML programs). Validation used an Apple M4 Max on macOS 27.0.1; other combinations are not qualified.
  • A host that implements the execution contract below. The Python reference runtime (encoder.py) does.
  • Disk for Core ML's compiled-model cache: a function's first load compiles it for the device, and later loads reuse the cache. Core ML keys the cache on the program's path, so moving the download compiles again. On the validation host a Neural Engine function's first load takes about 1 s and a later one about 0.14 s. The Neural Engine is shared by every process on the Mac and holds a limited number of loaded programs, so load the media functions when you need them. On the validation host the four GPU encoders' first loads take 5 to 9 s together, and later ones about 1 s together (up to 4 s).
  • Do not enable allowLowPrecisionAccumulationOnGPU: the GPU encoders were qualified with Core ML's default FP32 accumulation.

Download

hf download 1of2/bidirlm-omni-2.5b-coreml-w8-bundle --local-dir ./model

The download (about 2.8 GB) is the bundle: manifest.json at its root describes the compiled programs (BidirLMOmniLanguage.mlmodelc/, BidirLMOmniVision.mlmodelc/, BidirLMOmniAudio.mlmodelc/) and the host-side files in Resources/, and lists the SHA-256 of every file (checksums.files); verify them before loading, and keep the directory structure intact. Pin a Hub revision in production rather than following main.

Functions

Mode Functions
Neural Engine language_*: a head, 28 layer boundaries, attention over 512, 4,096, or 8,192 keys. vision_*: a head, 23 boundaries, a tail, the final patch merger, 2 DeepStack mergers, attention summaries over 512 to 4,096 keys. audio_*: the mel front end, a head, 23 boundaries, a tail, attention summaries over 512 to 4,096 keys, and merges of 2, 4, or 8 summaries
GPU language_gpu, vision_gpu, audio_front_gpu, audio_encoder_gpu (FP16 inputs and outputs)

A family's functions are in its own program: language_* in BidirLMOmniLanguage.mlmodelc, vision_* in BidirLMOmniVision.mlmodelc, and audio_* in BidirLMOmniAudio.mlmodelc. A family's GPU encoders read the same INT8 tensors as its Neural Engine functions, which its program stores once. manifest.json gives each Neural Engine stage's program and function name (streamed, streamedMedia). manifest.json gives each complete encoder's function, dtype, and declared lengths, with each family's program (complete).

Execution contract

  • Text. At most 8,192 tokens per message, including the chat template (<|im_start|>user\n ... <|im_end|>\n) and every expanded media placeholder. Longer input is an error, never truncated. Use the supplied tokenizer and processor.
  • Images. Resized within 65,536 to 1,048,576 pixels, 16-pixel patches merged 2 by 2, so an image becomes 64 to 1,024 tokens. The learned position table (vision_pos_embed.f32, 48 by 48) is interpolated on the host. Image tokens take three-axis (time, row, column) rotary positions, interleaved across frequency pairs; text after an image continues from the largest position used.
  • Audio. Mono, 16 kHz, 128 mel bins (window 400, hop 160): 25 audio tokens per 2 seconds, so a message holds at most about 10 minutes of audio.
  • Neural Engine staging. The language encoder runs a head per 512-token chunk, then one boundary per layer, with attention over every key of the message in between (keys in 4,096-key blocks whose softmax summaries merge exactly). The vision and audio towers run the same way over each item. Vision also returns DeepStack rows, added after the first two language layers at the image positions. Token-embedding lookup (token_embeddings.f16), rotary tables (cos.f16, sin.f16, inv_freq.f32), placeholder replacement, and pooling are host work.
  • GPU. Each complete encoder takes a whole item, padded to the next length it declares (Core ML enumerated shapes: 64 to 8,192 tokens for language and audio, 256 to 4,096 patches for vision), with an additive attention mask (bias: 0 for real rows, -30000 for padding and masked tokens). Language takes the embedding rows (token_embeddings.f16, with image and audio features in place), rotary tables computed on the host from inv_freq.f32, and two DeepStack row sets, and returns the final normalized hidden states; pooling is host work. The audio front end takes 20 mel chunks per call; the host gathers each chunk's tokens for the audio encoder.

Selecting .cpuAndNeuralEngine does not by itself guarantee Neural Engine execution. The staged functions were audited with Core ML compute plans on the validation host; check the plan on your own device if you need the guarantee. The GPU encoders were audited with Core ML compute plans on the validation host (every operation preferred and supported on the GPU).

Quality

Cosine similarity to the FP32 source model, through the reference runtime, per mode (release gate: 0.995):

Input Neural Engine GPU INT8 floor
Text, 97 tokens 0.99912 0.99912 0.99912
Text, 513 tokens, three masked 0.99872 0.99872 0.99872
Prose and code, 4,500 tokens 0.99890 0.99895 0.99895
Prose and code, 8,192 tokens 0.99848 0.99863 0.99863
Repeated sentence, 8,192 tokens 0.99760 0.99779 0.99779
Image (288 by 416) and text, 130 tokens 0.99854 0.99866 0.99864
Image (1,024 by 1,024), 1,031 tokens 0.99906 0.99906 0.99906
Audio (6 s) and text, 88 tokens 0.99837 0.99836 0.99836
Audio (39 s), 495 tokens 0.99847 0.99850 0.99850
Image, audio, and text, 113 tokens 0.99854 0.99854 0.99854
Two images and text, 191 tokens 0.99883 0.99887 0.99887

INT8 floor: the FP32 source model given the same INT8 weights, against the FP32 source. Each mode's agreement with that model, at least, over every embedding above: Neural Engine 0.999868, GPU 0.999999.

GPU embeddings agree with the Neural Engine's to at least 0.99987 on every case, so one index can hold vectors from either mode.

Inputs of one token repeated 8,191 times (tokens 334 and 9370) reach 0.99425 and 0.99485 on the Neural Engine, below the gate; the INT8 floor itself is 0.99432 and 0.99497, and the programs agree with it to 0.999927 and 0.999996, so the gap is the weight quantization, not the programs. The GPU reaches the same floor.

Pooling hides one property of the INT8 weights: the image and audio towers' feature rows, before the language encoder averages them, agree less closely with the FP32 source than the embeddings do. qualification.json lists each tower case's features against the FP32 source and against an FP32 model given the same INT8 weights; the programs add almost nothing to that floor. The Neural Engine functions use the tanh form of GELU everywhere, within 4.7e-4 of the exact form the source uses in its audio tower and patch mergers. The GPU encoders use the source's own forms of GELU.

qualification.json records every check: the gates, each case's measurements, the software versions, and the host.

Python reference runtime

encoder.py runs every function of the release. Python 3.12 and requirements.txt:

from PIL import Image
from encoder import Encoder, IMAGE, AUDIO

model = Encoder("./model")                  # Neural Engine; compute="gpu" for the GPU
text_vector = model.embed("How do I cool an overheating laptop?")
image_vector = model.embed(images=[Image.open("example.jpg")])
mixed_vector = model.embed("Compare this image " + IMAGE, images=[Image.open("example.jpg")])
# audio: lists of mono float32 waveforms at 16 kHz, one AUDIO placeholder each

Known issue. coremltools 9.0, the latest release, hands NumPy inputs to Core ML without copying and lets Core ML release them on a private queue without the Python GIL. Python 3.12 and later can then segfault some time after a prediction returns. The fix is merged upstream (apple/coremltools#2829) but not yet released; this release was qualified with 9.0 plus that change. Until a release includes it, install a coremltools build that does.

Insert one IMAGE or AUDIO placeholder per supplied item, in order. Do not use one Encoder from several threads at once; release("vision") and release("audio") unload a family's functions. The processor code in Resources/tokenizer/ runs locally with trust_remote_code=True; review it before use.

Files

  • manifest.json: the execution contract, the qualification seal, and the SHA-256 of every other file except this card, LICENSE, and the runtime.
  • BidirLMOmniLanguage.mlmodelc/, BidirLMOmniVision.mlmodelc/, BidirLMOmniAudio.mlmodelc/: the compiled programs, one per family.
  • Resources/: host-side tables, tokenizer, and processor.
  • streamed.json: the language stages.
  • qualification.json: the release's qualification evidence.
  • encoder.py, requirements.txt: the Python reference runtime.
  • LICENSE (Apache 2.0), THIRD_PARTY.md.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 1of2/bidirlm-omni-2.5b-coreml-w8-bundle

Unable to build the model tree, the base model loops to the model itself. Learn more.