BidirLM Omni 2.5B, Core ML W8A16, Neural Engine and GPU, 8K context
Compiled Core ML programs for text, still-image, audio, and mixed-message embeddings, one per family (language, vision, audio), with a path for the Mac's Neural Engine and one for its GPU over each family's one set of INT8 weights:
- Neural Engine: 97 staged functions; every operation was audited and actual Neural Engine execution was traced when the release was built.
- GPU: one complete encoder per family (language, vision, audio front end and encoder), a whole item per prediction. FP16 matrix work with an FP32 residual stream; the faster mode for long inputs.
A message holds up to 8,192 tokens, chat template and media included, in either mode. Both modes embed into the same space: the GPU's vectors are gated against the Neural Engine's, so one index can hold vectors from either. There is no CPU mode.
Every message produces the same kind of vector: 2,048 dimensions, L2-normalized, masked mean pooling over all of the message's tokens (chat-template tokens included). There are no query or document prompts; queries and documents are encoded identically.
Source: BidirLM-Omni-2.5B-Embedding,
revision 447a6e31be61b84443144afda21374339ce408e6. Projection weights use per-channel INT8 with
FP16 activations; the token-embedding table stays FP16.
Releases
The same functions are published in these layouts, and every one embeds into the same space (the
manifest's spaceID, ...:w8a16:ane-8k-v1), so one index can hold vectors from any of them:
| Repository | Modes | Programs |
|---|---|---|
| 1of2/bidirlm-omni-2.5b-coreml-w8 | Neural Engine and GPU | one program for every function (revision 57509034) |
| 1of2/bidirlm-omni-2.5b-coreml-w8-bundle (this repository) | Neural Engine and GPU | one program per family |
| 1of2/bidirlm-omni-2.5b-coreml-w8-ane | Neural Engine | one program per family |
| 1of2/bidirlm-omni-2.5b-coreml-w8-gpu | GPU | one program per family |
One program per family loads faster: Core ML reads a function's whole program each time it loads
one. Revisions of 1of2/bidirlm-omni-2.5b-coreml-w8 up to c228769d56fad334c82d726c31f184122f04f5f9
held a withdrawn release (r4, 32,768 tokens) that embeds into the source model's space too but is
not bit-identical: re-embed rather than mix its vectors into an index of these.
Requirements
- An Apple-silicon Mac with macOS 15 or later (multifunction Core ML programs). Validation used an Apple M4 Max on macOS 27.0.1; other combinations are not qualified.
- A host that implements the execution contract below. The Python reference runtime
(
encoder.py) does. - Disk for Core ML's compiled-model cache: a function's first load compiles it for the device, and later loads reuse the cache. Core ML keys the cache on the program's path, so moving the download compiles again. On the validation host a Neural Engine function's first load takes about 1 s and a later one about 0.14 s. The Neural Engine is shared by every process on the Mac and holds a limited number of loaded programs, so load the media functions when you need them. On the validation host the four GPU encoders' first loads take 5 to 9 s together, and later ones about 1 s together (up to 4 s).
- Do not enable
allowLowPrecisionAccumulationOnGPU: the GPU encoders were qualified with Core ML's default FP32 accumulation.
Download
hf download 1of2/bidirlm-omni-2.5b-coreml-w8-bundle --local-dir ./model
The download (about 2.8 GB) is the bundle: manifest.json at its root describes the compiled
programs (BidirLMOmniLanguage.mlmodelc/, BidirLMOmniVision.mlmodelc/,
BidirLMOmniAudio.mlmodelc/) and the host-side files in Resources/, and lists the SHA-256 of
every file (checksums.files); verify them before loading, and keep the directory structure
intact. Pin a Hub revision in production rather than following main.
Functions
| Mode | Functions |
|---|---|
| Neural Engine | language_*: a head, 28 layer boundaries, attention over 512, 4,096, or 8,192 keys. vision_*: a head, 23 boundaries, a tail, the final patch merger, 2 DeepStack mergers, attention summaries over 512 to 4,096 keys. audio_*: the mel front end, a head, 23 boundaries, a tail, attention summaries over 512 to 4,096 keys, and merges of 2, 4, or 8 summaries |
| GPU | language_gpu, vision_gpu, audio_front_gpu, audio_encoder_gpu (FP16 inputs and outputs) |
A family's functions are in its own program: language_* in BidirLMOmniLanguage.mlmodelc,
vision_* in BidirLMOmniVision.mlmodelc, and audio_* in BidirLMOmniAudio.mlmodelc.
A family's GPU encoders read the same INT8 tensors as its Neural Engine functions, which its
program stores once.
manifest.json gives each Neural Engine stage's program and function name (streamed,
streamedMedia).
manifest.json gives each complete encoder's function, dtype, and declared lengths, with each
family's program (complete).
Execution contract
- Text. At most 8,192 tokens per message, including the chat template
(
<|im_start|>user\n...<|im_end|>\n) and every expanded media placeholder. Longer input is an error, never truncated. Use the supplied tokenizer and processor. - Images. Resized within 65,536 to 1,048,576 pixels, 16-pixel patches merged 2 by 2, so an image
becomes 64 to 1,024 tokens. The learned position table (
vision_pos_embed.f32, 48 by 48) is interpolated on the host. Image tokens take three-axis (time, row, column) rotary positions, interleaved across frequency pairs; text after an image continues from the largest position used. - Audio. Mono, 16 kHz, 128 mel bins (window 400, hop 160): 25 audio tokens per 2 seconds, so a message holds at most about 10 minutes of audio.
- Neural Engine staging. The language encoder runs a head per 512-token chunk, then one boundary per layer,
with attention over every key of the message in between (keys in 4,096-key blocks whose softmax
summaries merge exactly). The vision and audio towers run the same way over each item. Vision
also returns DeepStack rows, added after the first two language layers at the image positions.
Token-embedding lookup (
token_embeddings.f16), rotary tables (cos.f16,sin.f16,inv_freq.f32), placeholder replacement, and pooling are host work. - GPU. Each complete encoder takes a whole item, padded to the next length it declares
(Core ML enumerated shapes: 64 to 8,192 tokens for language and audio, 256 to 4,096 patches for
vision), with an additive attention mask (
bias: 0 for real rows, -30000 for padding and masked tokens). Language takes the embedding rows (token_embeddings.f16, with image and audio features in place), rotary tables computed on the host frominv_freq.f32, and two DeepStack row sets, and returns the final normalized hidden states; pooling is host work. The audio front end takes 20 mel chunks per call; the host gathers each chunk's tokens for the audio encoder.
Selecting .cpuAndNeuralEngine does not by itself guarantee Neural Engine execution. The staged
functions were audited with Core ML compute plans on the validation host; check the plan on your
own device if you need the guarantee.
The GPU encoders were audited with Core ML compute plans on the validation host (every operation
preferred and supported on the GPU).
Quality
Cosine similarity to the FP32 source model, through the reference runtime, per mode (release gate: 0.995):
| Input | Neural Engine | GPU | INT8 floor |
|---|---|---|---|
| Text, 97 tokens | 0.99912 | 0.99912 | 0.99912 |
| Text, 513 tokens, three masked | 0.99872 | 0.99872 | 0.99872 |
| Prose and code, 4,500 tokens | 0.99890 | 0.99895 | 0.99895 |
| Prose and code, 8,192 tokens | 0.99848 | 0.99863 | 0.99863 |
| Repeated sentence, 8,192 tokens | 0.99760 | 0.99779 | 0.99779 |
| Image (288 by 416) and text, 130 tokens | 0.99854 | 0.99866 | 0.99864 |
| Image (1,024 by 1,024), 1,031 tokens | 0.99906 | 0.99906 | 0.99906 |
| Audio (6 s) and text, 88 tokens | 0.99837 | 0.99836 | 0.99836 |
| Audio (39 s), 495 tokens | 0.99847 | 0.99850 | 0.99850 |
| Image, audio, and text, 113 tokens | 0.99854 | 0.99854 | 0.99854 |
| Two images and text, 191 tokens | 0.99883 | 0.99887 | 0.99887 |
INT8 floor: the FP32 source model given the same INT8 weights, against the FP32 source. Each mode's agreement with that model, at least, over every embedding above: Neural Engine 0.999868, GPU 0.999999.
GPU embeddings agree with the Neural Engine's to at least 0.99987 on every case, so one index can hold vectors from either mode.
Inputs of one token repeated 8,191 times (tokens 334 and 9370) reach 0.99425 and 0.99485 on the Neural Engine, below the gate; the INT8 floor itself is 0.99432 and 0.99497, and the programs agree with it to 0.999927 and 0.999996, so the gap is the weight quantization, not the programs. The GPU reaches the same floor.
Pooling hides one property of the INT8 weights: the image and audio towers' feature rows, before
the language encoder averages them, agree less closely with the FP32 source than the embeddings
do. qualification.json lists each tower case's features against the FP32 source and against an
FP32 model given the same INT8 weights; the programs add almost nothing to that floor.
The Neural Engine functions use the tanh form of GELU everywhere, within 4.7e-4 of the exact form
the source uses in its audio tower and patch mergers.
The GPU encoders use the source's own forms of GELU.
qualification.json records every check: the gates, each case's measurements, the software
versions, and the host.
Python reference runtime
encoder.py runs every function of the release. Python 3.12 and requirements.txt:
from PIL import Image
from encoder import Encoder, IMAGE, AUDIO
model = Encoder("./model") # Neural Engine; compute="gpu" for the GPU
text_vector = model.embed("How do I cool an overheating laptop?")
image_vector = model.embed(images=[Image.open("example.jpg")])
mixed_vector = model.embed("Compare this image " + IMAGE, images=[Image.open("example.jpg")])
# audio: lists of mono float32 waveforms at 16 kHz, one AUDIO placeholder each
Known issue. coremltools 9.0, the latest release, hands NumPy inputs to Core ML without copying and lets Core ML release them on a private queue without the Python GIL. Python 3.12 and later can then segfault some time after a prediction returns. The fix is merged upstream (apple/coremltools#2829) but not yet released; this release was qualified with 9.0 plus that change. Until a release includes it, install a coremltools build that does.
Insert one IMAGE or AUDIO placeholder per supplied item, in order. Do not use one Encoder
from several threads at once; release("vision") and release("audio") unload a family's
functions. The processor code in Resources/tokenizer/ runs locally with
trust_remote_code=True; review it before use.
Files
manifest.json: the execution contract, the qualification seal, and the SHA-256 of every other file except this card,LICENSE, and the runtime.BidirLMOmniLanguage.mlmodelc/,BidirLMOmniVision.mlmodelc/,BidirLMOmniAudio.mlmodelc/: the compiled programs, one per family.Resources/: host-side tables, tokenizer, and processor.streamed.json: the language stages.qualification.json: the release's qualification evidence.encoder.py,requirements.txt: the Python reference runtime.LICENSE(Apache 2.0),THIRD_PARTY.md.
Model tree for 1of2/bidirlm-omni-2.5b-coreml-w8-bundle
Unable to build the model tree, the base model loops to the model itself. Learn more.