Commit Graph

  • 240921fd96 fix: omit array parsing mxyng/omit-array Michael Yang 2025-05-15 16:26:02 -0700
  • 27da2cddc5
    Fix lingering Q4_0 help reference (#10720) Daniel Hiltgen 2025-05-15 16:33:23 -0700
  • feb8923ada
    cmd: add ellipses to truncated show metadata (#10717) Bruce MacDonald 2025-05-15 15:45:52 -0700
  • 2ab70e82d0 remove rebase debug parth/tool-prefix-temp ParthSareen 2025-05-15 14:57:42 -0700
  • 717fa7a44a Add sentinel errors, remove redundant calls ParthSareen 2025-05-15 14:23:14 -0700
  • fe623c2cf4 ollamarunner: Multi-modal worst case graph Jesse Gross 2025-04-07 13:59:11 -0700
  • 3c14461d5d ollamarunner: Separate text and multimodal graphs Jesse Gross 2025-05-05 13:32:11 -0700
  • 499ae7311f ollamarunner: Base cached tokens on current prompt Jesse Gross 2025-05-09 16:51:47 -0700
  • ef202789fa
    fix pixel values padding (#10718) Michael Yang 2025-05-15 13:44:44 -0700
  • 55760195e6
    fix mllama conversion (#10716) Michael Yang 2025-05-15 12:15:01 -0700
  • bd68d3ae50
    ggml: update qwen25vl vision size estimate (#10711) v0.7.0 Bruce MacDonald 2025-05-14 16:42:30 -0700
  • 53f7946fb6 add tests, organize, comments ParthSareen 2025-05-13 16:51:55 -0700
  • ff80718e9c
    fix crash in old clients with quantization progress (#10710) Daniel Hiltgen 2025-05-14 14:54:18 -0700
  • 5a2cd7b48a runner: add test for unicode token processing brucemacd/runner-test Bruce MacDonald 2025-05-14 11:29:11 -0700
  • 0aa8b371dd
    model: add Qwen2.5-VL support (#10385) v0.7.0-rc1 Bruce MacDonald 2025-05-13 20:58:02 -0700
  • bc83789be9 tools package and utils ParthSareen 2025-05-12 18:02:18 -0700
  • 4059b8db01 renaming and splitting stuff up ParthSareen 2025-05-12 14:07:59 -0700
  • b8b9c0c7cf checkpoint ParthSareen 2025-05-09 17:05:16 -0700
  • 779547fcde checkpoint - cleanup still left, functionality setup ParthSareen 2025-05-08 18:48:44 -0700
  • 6cb7494061 checkpoint for new parser ParthSareen 2025-05-07 19:35:11 -0700
  • a44734b030 add new parser, tests, and templates ParthSareen 2025-05-07 15:50:51 -0700
  • b5a982ecb0 wip ParthSareen 2025-05-06 18:29:06 -0700
  • 516a540df7 jsonv2 decoder ParthSareen 2025-05-05 17:25:35 -0700
  • 7f2f996cd6 server/routes: catch when JSON tool was used ParthSareen 2025-05-02 14:16:39 -0700
  • 610054a234 model: support tools streaming and improve parsing ParthSareen 2025-04-25 16:35:16 -0700
  • 23125648b8
    chore: update mllama to use ollama engine (#10637) Michael Yang 2025-05-13 17:36:02 -0700
  • 0478d440f0
    Fixed over vram allcation dure to small initial layer sizes. tej 2025-05-13 18:42:39 -0500
  • 8cc33f4c2b
    llama: fix memory leak for grammar (#10696) Parth Sareen 2025-05-13 15:39:27 -0700
  • f46df4e5d2
    llama: fix defrag patch to defragment when no slots are available (#10695) Jeffrey Morgan 2025-05-13 14:02:08 -0700
  • c6bcdc4223
    Revert "remove cuda v11 (#10569)" (#10692) Daniel Hiltgen 2025-05-13 13:12:54 -0700
  • 4b903f088a
    llama: fix crash on snowflake embedding model (#10690) Jeffrey Morgan 2025-05-13 13:11:11 -0700
  • c7f4ae7b9c
    server: add webp image input support (#10653) Jeffrey Morgan 2025-05-12 20:41:42 -0700
  • 5c76074f66 wip jmorganca/qwen25vl jmorganca 2025-05-12 19:15:42 -0700
  • 526b2ed102
    fix vocabulary (#10679) Michael Yang 2025-05-12 17:29:46 -0700
  • a7240c6d63
    models: remove unused qwen2vl processing (#10677) Bruce MacDonald 2025-05-12 16:08:42 -0700
  • 9d6df90805
    Follow up to #10363 (#10647) v0.7.0-rc0 Daniel Hiltgen 2025-05-12 15:23:31 -0700
  • 18d52686de
    Update model/models/qwen25vl/model_vision.go Bruce MacDonald 2025-05-12 14:16:46 -0700
  • 2d2eb5903d use with pattern for rope Bruce MacDonald 2025-05-12 14:14:03 -0700
  • 533f4c41bd add eot Bruce MacDonald 2025-05-12 14:03:37 -0700
  • 31b2c06393 Update 0007-add-unpad-operator.patch Bruce MacDonald 2025-05-12 13:51:46 -0700
  • 4ae23deb50 Revert "Update 0007-add-unpad-operator.patch" Bruce MacDonald 2025-05-12 13:43:53 -0700
  • 5d3da85a16 remove out of date comments Bruce MacDonald 2025-05-12 13:26:26 -0700
  • 8b64b456c1 Update 0007-add-unpad-operator.patch Bruce MacDonald 2025-05-12 13:24:55 -0700
  • 684f0d9291 set default values for vision model in config Bruce MacDonald 2025-05-12 12:10:15 -0700
  • 3308bff137 add i32 copy and argsort for cuda jmorganca 2025-05-10 17:34:53 -0700
  • bf1929a3bc Delete 0017-add-ollama-vocab-for-grammar-support.patch Bruce MacDonald 2025-05-09 15:11:17 -0700
  • 1a2c413225 move mask Bruce MacDonald 2025-05-09 14:35:10 -0700
  • 57279f89a2 calculate block mask once, rather than in attention Bruce MacDonald 2025-05-09 13:58:47 -0700
  • 9ceee25d8b chunk vision outputs Bruce MacDonald 2025-05-08 16:31:27 -0700
  • 661bf04696 add picture prefix Bruce MacDonald 2025-05-08 11:51:00 -0700
  • 2521a55ae6 fixes after rebase Bruce MacDonald 2025-05-08 09:30:18 -0700
  • 32948ec952 increase rope base Bruce MacDonald 2025-05-07 15:56:42 -0700
  • 9876c8453a update exported functions for tests Bruce MacDonald 2025-05-07 15:35:29 -0700
  • 919b3d6e21 require new engine for qwen25vl arch Bruce MacDonald 2025-05-02 15:54:15 -0700
  • 16b13e0cfc Revert "ropeTheta should be 1e5" Bruce MacDonald 2025-05-02 15:42:35 -0700
  • 75441c56f3 add comment explaining rope theta Bruce MacDonald 2025-05-02 15:16:22 -0700
  • 45f96e898d ropeTheta should be 1e5 Bruce MacDonald 2025-05-02 15:09:37 -0700
  • 7c555d394c simplify patch creation Bruce MacDonald 2025-05-01 14:36:07 -0700
  • 39ee6d2bd0 ranges for lint Bruce MacDonald 2025-05-01 14:06:30 -0700
  • 47705b5168 simplify rope changes Bruce MacDonald 2025-05-01 13:47:55 -0700
  • 698a92aa4a reverse window Michael Yang 2025-05-01 13:45:54 -0700
  • 150c499cae use silu Michael Yang 2025-05-01 12:49:02 -0700
  • f1257a7de4 update vision rope theta default Michael Yang 2025-05-01 12:37:21 -0700
  • b68af0370f move sdpa to model forward pass Bruce MacDonald 2025-05-01 11:51:32 -0700
  • ca981c8a49 full attn block indexes should be []int32 Bruce MacDonald 2025-05-01 11:21:56 -0700
  • b3da8a319e Update model_vision.go Bruce MacDonald 2025-05-01 11:06:06 -0700
  • 359e1d5b19 full attention layers Bruce MacDonald 2025-05-01 10:59:46 -0700
  • bde6b46ce9 fix padding Michael Yang 2025-04-30 17:17:58 -0700
  • ff1f74534b block attention Bruce MacDonald 2025-04-30 16:43:37 -0700
  • 104f802df1 remove todos Bruce MacDonald 2025-04-29 16:57:41 -0700
  • eed0ac2948 clean up vision model forward pass Bruce MacDonald 2025-04-29 16:54:48 -0700
  • fcfad744ff fix patch merger Bruce MacDonald 2025-04-29 15:43:10 -0700
  • fb3c16f2a2 window index Michael Yang 2025-04-29 14:46:23 -0700
  • ee869f35e4 fix image processing Michael Yang 2025-04-29 11:31:49 -0700
  • ff5d1a3dc0 duplicate input embeddings Michael Yang 2025-04-29 10:09:44 -0700
  • 88b231f903 use maxgridsize Michael Yang 2025-04-29 09:58:17 -0700
  • 7e920c8d75 fix: patch merger and convert Michael Yang 2025-04-28 13:59:54 -0700
  • dd8c619fba fixes after rebase Bruce MacDonald 2025-04-28 14:19:26 -0700
  • 2af76d0e7a default to 32 for vision block count Bruce MacDonald 2025-04-28 13:09:27 -0700
  • 8d901825f0 reshape cos and sin Bruce MacDonald 2025-04-28 11:14:12 -0700
  • 04936b719f Update model_vision.go Bruce MacDonald 2025-04-28 09:41:04 -0700
  • 0f0136d419 simplify by doing operations in Go rather than with tensors Bruce MacDonald 2025-04-24 17:48:00 -0700
  • 80498f76de fix build Bruce MacDonald 2025-04-23 16:22:22 -0700
  • f8b48aa784 Delete model_external_test.go Bruce MacDonald 2025-04-23 16:13:08 -0700
  • 5ff0d538b0 wip: implementing rope Bruce MacDonald 2025-04-21 18:50:36 -0700
  • eedc969c35 grid refactor Bruce MacDonald 2025-04-21 15:06:13 -0700
  • 963531215e update convert Bruce MacDonald 2025-04-21 13:58:26 -0700
  • 3fe090f447 get patch embedding vals from config Bruce MacDonald 2025-04-21 12:04:46 -0700
  • 1704072746 patch embeddings Bruce MacDonald 2025-04-21 09:43:56 -0700
  • c1f9bcb4dd restructure Bruce MacDonald 2025-04-02 10:41:51 -0700
  • 198b1e6db9 text model forward pass Bruce MacDonald 2025-04-01 14:09:41 -0700
  • 51ad65f831 ml: structured rope config to allow specifying context len Bruce MacDonald 2025-04-01 14:03:48 -0700
  • 0cefd46f23
    llama: update to commit de4c07f93 (#10655) Jeffrey Morgan 2025-05-12 12:17:26 -0700
  • ad035ad595
    convert: quantize from safetensors needs kv (#10675) Bruce MacDonald 2025-05-12 12:04:20 -0700
  • f95a1f2bef
    feat: add trace log level (#10650) Michael Yang 2025-05-12 11:43:00 -0700
  • 82a9e9462a
    readme: add UnityCodeLama to community integrations (#10665) HardCodeDev 2025-05-12 00:44:51 +0400
  • 76724e2f29
    readme: add OllamaPlusPlus C++ library to community integrations (#10664) HardCodeDev 2025-05-12 00:40:41 +0400
  • ecf14a220f
    llama: allocate grammar buffer based on schema length (#10649) frob 2025-05-10 20:57:30 +0200
  • 69ce44b33c
    envconfig: Remove no longer supported max vram var (#10623) frob 2025-05-10 20:31:04 +0200
  • 5969674cf1
    feat: add threshold to dump options (#10639) Michael Yang 2025-05-10 11:27:15 -0700