Commit Graph

  • 837379a94c
    discovery: fix cudart driver version (#11614) main Daniel Hiltgen 2025-08-13 15:43:33 -0700
  • a24f90604f
    int: adjust a few models for integration tests (#11872) Daniel Hiltgen 2025-08-13 15:42:36 -0700
  • dc5a645434
    cuda: leverage JIT for smaller footprint (#11635) Daniel Hiltgen 2025-08-13 15:42:16 -0700
  • e5cc4528ad clean up patches gpt-oss-bump Michael Yang 2025-08-13 15:23:49 -0700
  • 8b4d65cd17 llm: New memory management jessegross/bump-memory Jesse Gross 2025-05-29 12:21:48 -0700
  • cb81ea7772 update patch instructions Michael Yang 2025-08-13 14:39:21 -0700
  • 12226de6d1 Merge branch 'main' of github.com:ollama/ollama into gpt-oss-bump Michael Yang 2025-08-13 14:15:28 -0700
  • a495b8e128 llm: New memory management jessegross/memory Jesse Gross 2025-05-29 12:21:48 -0700
  • b5c10e43c4 clean up vendor Michael Yang 2025-08-13 11:54:27 -0700
  • 752e41c26e refactor cpu quants Michael Yang 2025-08-13 11:50:18 -0700
  • e149575110 add back i32->i32 copy Michael Yang 2025-08-13 11:40:31 -0700
  • bb71654ebe chore: fix some inconsistent function name in comment youzichuan 2025-08-13 16:22:45 +0800
  • a343ae53a4 ggml: Use ordinal IDs for AMD GPUs on Linux when UUID is unavailable Jesse Gross 2025-08-11 17:01:07 -0700
  • 478f47ef8b server: add debug option for printing out prompt instead of calling model drifkin/print-template Devon Rifkin 2025-08-12 16:55:54 -0700
  • 908f199ab2 remap values only if source/target are different Michael Yang 2025-08-11 13:47:07 -0700
  • 0296f325ff remove redundant test Michael Yang 2025-08-09 12:07:01 -0700
  • 274759bb9a fix lint Michael Yang 2025-08-09 11:58:18 -0700
  • 5f4474fa1c remove debug messages Michael Yang 2025-08-09 11:56:49 -0700
  • 3d8e613845 fast attention Michael Yang 2025-08-08 15:12:18 -0700
  • a9e4c87686 fix nested alt tags Michael Yang 2025-08-07 03:42:02 -0700
  • 8233ae3ecd split qkv, gate_up Michael Yang 2025-08-07 02:19:56 -0700
  • 6d86ad2536 add ids Michael Yang 2025-08-06 19:57:16 -0700
  • 91473866e7 openai swiglu Michael Yang 2025-08-06 19:33:42 -0700
  • b56f002e7a reshape earlier Michael Yang 2025-08-06 19:26:47 -0700
  • 2e84c8bcd6 buffer the conversion better Daniel Hiltgen 2025-08-07 11:44:20 -0700
  • afdffa599a convert mlp bf16 to f32 Daniel Hiltgen 2025-08-07 11:04:15 -0700
  • 6df2806211 Convert tensors at load time Daniel Hiltgen 2025-08-07 09:39:30 -0700
  • 86ef7f7e15 fix windows build error Daniel Hiltgen 2025-08-06 12:38:49 -0700
  • 559d8dd23a bump Daniel Hiltgen 2025-08-06 11:35:15 -0700
  • 18b44c3186 unwind mxfp4 patch Daniel Hiltgen 2025-08-06 09:48:39 -0700
  • 930f16920d fix: Sync ggml-cuda.cu after keeping both style cuda graph fixes for gemma3n Gabe Goodhart 2025-07-30 14:51:17 -0400
  • 441955e586 fix: Update 0020 CUDA Graphs for gemma3n to keep both llama.cpp and ollama fixes Gabe Goodhart 2025-07-30 14:50:42 -0400
  • 8208b8f3e7 Revert "fix: Remove Gemma3n CUDA Graphs patch" Gabe Goodhart 2025-07-30 14:34:29 -0400
  • 20a5fe5a0c fix: Remove unused vendored code for chat template parsing Gabe Goodhart 2025-07-30 14:11:34 -0400
  • 612c0edd03 fix: Remove unnecessary additions in the rsync-filter Gabe Goodhart 2025-07-30 14:06:44 -0400
  • aae20db57b build: Remove unnecessary CFLAGS definitions in cpu.go Gabe Goodhart 2025-07-30 14:02:42 -0400
  • c880be3f44 feat: Sync llama.cpp / ggml after latest bump Gabe Goodhart 2025-07-30 13:59:11 -0400
  • fd66dc2d38 fix: Remove Gemma3n CUDA Graphs patch Gabe Goodhart 2025-07-30 13:55:21 -0400
  • dedfeb4a5b fix: Fix Solar and argsort/copy patches after bump Gabe Goodhart 2025-07-30 13:54:38 -0400
  • 1fca926e3d feat: Bump to 41e78c in the makefile Gabe Goodhart 2025-07-30 13:54:03 -0400
  • c09dbc6038 fix: Re-number patches after merge with `main` Gabe Goodhart 2025-07-30 13:02:08 -0400
  • 0c2c1bcf7c fix: Handle multi-chunk image encodings from mtmd Gabe Goodhart 2025-07-28 11:38:39 -0400
  • a5ef6c67fd feat: Sync llama.cpp Gabe Goodhart 2025-07-15 14:50:01 -0600
  • 9bb45f633e feat: Bump llama.cpp to 4a4f42 Gabe Goodhart 2025-07-15 14:49:15 -0600
  • d6ab2adea0 fix: Sync patch changes for ggml-cpu.c Gabe Goodhart 2025-07-11 16:01:15 -0600
  • 3580ac02ab fix: Add a patch to avoid power throttling API on non-msvc windows builds Gabe Goodhart 2025-07-11 16:00:49 -0600
  • 341f07b1f0 build: Add top-level include for GNUINstallDirs in CMakeLists.txt Gabe Goodhart 2025-07-11 13:44:10 -0600
  • 4514aa6548 build: Include cmake/common.cmake in ggml sync Gabe Goodhart 2025-07-11 13:25:01 -0600
  • 9ce694b951 feat: Sync all patched code Gabe Goodhart 2025-07-11 11:44:18 -0600
  • a5c3f40676 fix: Add patch for GGML_VERSION and GGML_COMMIT constants Gabe Goodhart 2025-07-11 11:43:14 -0600
  • 94a7a57fb0 fix: Revert changes to ggml export GPU UUID patch Gabe Goodhart 2025-07-11 11:42:26 -0600
  • 87920eac1a feat: Bump back to the cenral repo and point at the latest master Gabe Goodhart 2025-07-11 10:43:22 -0600
  • accfa536ac fix: Update patches for bump Gabe Goodhart 2025-07-10 16:01:30 -0600
  • 58bd438c9f feat: Bump to the latest tip of the branch Gabe Goodhart 2025-07-10 16:01:14 -0600
  • 5794a9e0fc fix: Update patch 0015 for upstream implementation of uuid Gabe Goodhart 2025-07-10 14:33:12 -0600
  • f49e1c7e08 fix: Use c++17 and include vendor for go wrapper modules Gabe Goodhart 2025-06-27 17:17:45 -0600
  • 3a6688eab4 fix: Add sync'ed stb vendored header Gabe Goodhart 2025-06-27 17:17:23 -0600
  • 7dde263642 fix: Add missing stb to llama.cpp rsync-filter Gabe Goodhart 2025-06-27 17:16:58 -0600
  • 561926a25a fix: Apply patch for mtmd_text_input Gabe Goodhart 2025-06-27 17:09:48 -0600
  • f413332d5a fix: Use mtmd_helper to correctly load the bitmap for the image Gabe Goodhart 2025-06-26 16:47:09 -0600
  • 8c06728134 fix: Fix support for arch-specific ggml-cpu source files with new arrangement Gabe Goodhart 2025-06-25 08:38:58 -0600
  • 497bed01d8 chore: Ignore *.patched in the patch directory Gabe Goodhart 2025-06-25 06:36:32 -0600
  • fba01b45b5 fix: Add patch for mtmd_input_text Gabe Goodhart 2025-06-25 06:35:18 -0600
  • 4a145edf74 fix: Update llama.go to use mtmd instead of clip/llava Gabe Goodhart 2025-06-24 17:48:31 -0600
  • 64d4dcaf02 fix: Add missing include in sampling_ext.cpp Gabe Goodhart 2025-06-24 17:47:56 -0600
  • a0249d4fff fix: Remove mtmd main cpp files Gabe Goodhart 2025-06-24 17:47:36 -0600
  • 1e02876646 fix: Narrow llama.cpp rsync-filter to not include mtmd main tool cpp files Gabe Goodhart 2025-06-24 17:46:54 -0600
  • 59d7586765 fix: Add ggml files missing from sync Gabe Goodhart 2025-06-27 17:06:05 -0600
  • 22d2205616 fix: Update ggml rsync-filter for new ggml-cpu/arch subdirs Gabe Goodhart 2025-06-24 17:44:52 -0600
  • 740a1aef00 fix: Add files missing from sync Gabe Goodhart 2025-06-27 17:05:25 -0600
  • 2d6aaf958f fix: Update rsync-filter for all moved/new/removed files Gabe Goodhart 2025-06-27 17:04:51 -0600
  • d206bc8c86 feat: Sync llama.cpp and ggml Gabe Goodhart 2025-06-27 17:01:24 -0600
  • 68a22ac1ed feat: Update all patches Gabe Goodhart 2025-06-27 16:57:05 -0600
  • 7b45a5520e TEMPORARY: Update the llama.cpp upstream to my fork's Granite Four branch Gabe Goodhart 2025-06-27 16:24:42 -0600
  • d0cf6c8281
    fix(openai): handle reasoning_effort (#11868) Michael Yang 2025-08-12 11:02:01 -0700
  • 8f4ec9ab28 discover: CPU supports flash attention Jesse Gross 2025-08-11 14:45:45 -0700
  • dbfd7bd027
    Merge pull request #11861 from ollama/drifkin/fix-parsing-error Devon Rifkin 2025-08-11 14:59:57 -0700
  • ee04dbba51 server: fix error when parsing bad harmony tool calls Devon Rifkin 2025-08-11 14:09:13 -0700
  • ea7657b54a
    sched: Add support for grouping GPUs (#10678) Daniel Andersen 2025-08-11 22:59:38 +0200
  • 2c776f0780
    CONTRIBUTING: Explicitly note docs:... as a good example (#11755) Michael Vorburger 2025-08-10 03:12:30 +0200
  • 0c3c32d24f add model benchmark mxyng/benchmark Michael Yang 2025-08-08 14:54:32 -0700
  • 79f6376f5b ggml: No-alloc mode Jesse Gross 2025-07-23 14:18:24 -0700
  • 756c78cfc7 ggml: Support closing backends Jesse Gross 2025-04-17 17:12:01 -0700
  • d7f4f788d1 ggml: Use GGML's typedef'ed pointer types Jesse Gross 2025-08-06 11:39:08 -0700
  • 114c3f2265
    tests: add integration coverage for oss-gpt (#11696) Daniel Hiltgen 2025-08-07 15:06:57 -0700
  • f2e9c9aff5 server: Reduce gpt-oss context length for small VRAM GPUs v0.11.4 Jesse Gross 2025-08-07 13:49:26 -0700
  • aa9d889522
    Merge pull request #11765 from ollama/drifkin/thinking-without-content v0.11.4-rc0 Devon Rifkin 2025-08-06 19:02:23 -0700
  • 735c41f9ca openai: always provide reasoning Devon Rifkin 2025-08-06 18:54:20 -0700
  • 223a619468
    Merge pull request #11761 from ollama/drifkin/openai-tool-names Devon Rifkin 2025-08-06 17:53:25 -0700
  • 759dd78dd6 openai: when converting role=tool messages, propagate the tool name Devon Rifkin 2025-08-06 17:00:24 -0700
  • 44bc36d063
    docs: update the faq (#11760) Patrick Devine 2025-08-06 16:55:57 -0700
  • 8f14e1f5f6
    Merge pull request #11759 from ollama/drifkin/oai-tool-calling Devon Rifkin 2025-08-06 16:11:31 -0700
  • 203c137810 openai: allow for content _and_ tool calls in the same message Devon Rifkin 2025-08-06 15:50:30 -0700
  • fa8be9e35c
    clean up debugging (#11756) Daniel Hiltgen 2025-08-06 13:31:22 -0700
  • 8a75e9ee15
    Update downloading to pulling in api.md (#11170) Gao feng 2025-08-07 02:33:09 +0800
  • 22ff54f018 update tests mxyng/16-bit Michael Yang 2025-08-05 20:43:38 -0700
  • 6779187517 drop float16 dependency Michael Yang 2025-08-05 20:06:33 -0700
  • 30f71416b9 drop bfloat16 dependency Michael Yang 2025-08-05 17:32:22 -0700
  • 4742e12c23
    docs: update turbo model name (#11707) v0.11.3 Parth Sareen 2025-08-05 17:29:08 -0700
  • 2d06977ade
    Merge pull request #11705 from ollama/drifkin/fn-schema v0.11.3-rc0 Devon Rifkin 2025-08-05 17:02:42 -0700