Add/update the quantized ONNX model files and README.md for Transformers.js v3 (#2)

- Add/update the quantized ONNX model files and README.md for Transformers.js v3 (fb5dbab37b4a3a352753bc1039dd8a202394c245)

Co-authored-by: Yuichiro Tachibana <[email protected]>

Files changed (6) hide show

README.md CHANGED Viewed

@@ -6,17 +6,20 @@ pipeline_tag: feature-extraction
 https://huggingface.co/jinaai/jina-embeddings-v2-base-en with ONNX weights to be compatible with Transformers.js.
 ## Usage with 🤗 Transformers.js
 ```js
-// npm i @xenova/transformers
-import { pipeline, cos_sim } from '@xenova/transformers';
 // Create feature extraction pipeline
-const extractor = await pipeline('feature-extraction', 'Xenova/jina-embeddings-v2-base-en',
-    { quantized: false } // Comment out this line to use the quantized version
-);
 // Generate embeddings
 const output = await extractor(
@@ -28,5 +31,4 @@ const output = await extractor(
 console.log(cos_sim(output[0].data, output[1].data));  // 0.9341313949712492 (unquantized) vs. 0.9022937687830741 (quantized)
 ```
 Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using [🤗 Optimum](https://huggingface.co/docs/optimum/index) and structuring your repo like this one (with ONNX weights located in a subfolder named `onnx`).

 https://huggingface.co/jinaai/jina-embeddings-v2-base-en with ONNX weights to be compatible with Transformers.js.
 ## Usage with 🤗 Transformers.js
+If you haven't already, you can install the [Transformers.js](https://huggingface.co/docs/transformers.js) JavaScript library from [NPM](https://www.npmjs.com/package/@huggingface/transformers) using:
+```bash
+npm i @huggingface/transformers
+```
 ```js
+import { pipeline, cos_sim } from '@huggingface/transformers';
 // Create feature extraction pipeline
+const extractor = await pipeline('feature-extraction', 'Xenova/jina-embeddings-v2-base-en', {
+    dtype: "fp32"  // Options: "fp32", "fp16", "q8", "q4"
+});
 // Generate embeddings
 const output = await extractor(
 console.log(cos_sim(output[0].data, output[1].data));  // 0.9341313949712492 (unquantized) vs. 0.9022937687830741 (quantized)
 ```
 Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using [🤗 Optimum](https://huggingface.co/docs/optimum/index) and structuring your repo like this one (with ONNX weights located in a subfolder named `onnx`).

onnx/model_bnb4.onnx ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:942109bb9c0654eca2aec01d23b7ad9f5d84ea4cc3622a059d8e29dd5df77386
+size 158113971

onnx/model_int8.onnx ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:5a01f42d114a6bb92228d1a47c4cccef79cfd931791a499134db80e177a3a93d
+size 137390704

onnx/model_q4.onnx ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:180bd960d42efb0f4671e34776461ff54f19f4b9471dfa07162b3ff5e5c3ae80
+size 165191319

onnx/model_q4f16.onnx ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:10f89500a7dfd28b4aeaf87b9ee03d9c8c37befeb231b27615be63e8354df3b5
+size 111050734

onnx/model_uint8.onnx ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:cdc0b9730a2cbae3edee4ff1869a45aeb6874e0a0d505f1b509b4e30964ace2c
+size 137390742