4153 Commits (parth/python-tools-calling)
 

Author SHA1 Message Date
Blake Mizerany 3519dd1c6e
server/internal/client/ollama: hold DiskCache on Registry (#9463) 1 year ago
Jeffrey Morgan e41c4cbea7
build: install ccache manually in Dockerfile (#9464) 1 year ago
Blake Mizerany ee048b76d4
server/internal/client/ollama: handle extended names in client/ollama (#9454) 1 year ago
Soulter af68d60a58
readme: add AstrBot to community integrations (#9442) 1 year ago
Jesse Gross 21aa666a1e ml: Enable support for flash attention 1 year ago
Jesse Gross ee141cc821 ml: Empty tensor constructor for tensors 1 year ago
Jesse Gross 55e5776c44 ggml-backend: Store parent backend as part of tensor 1 year ago
Jesse Gross 854a9195f3 attention: Remove unnecessary contiguous operations 1 year ago
Jeffrey Morgan 96a97adf9b
build: use correct GGML_HIP_NO_VMM compiler definition for ggml-hip (#9451) 1 year ago
Jeffrey Morgan e75c6126e9
build: set GGML_CUDA_NO_VMM for ggml-hip target (#9449) 1 year ago
Blake Mizerany cda6f5c66c
server/internal/internal/names: validate names (#9400) 1 year ago
Bruce MacDonald bebb6823c0
server: validate local path on safetensor create (#9379) 1 year ago
Michael Yang 31e472baa4 runner: defer context cancel 1 year ago
Michael Yang 657685e85d fix: replace deprecated functions 1 year ago
Jeffrey Morgan a14912858e
build: add compute capability 12.0 to CUDA 12 preset (#9426) 1 year ago
Blake Mizerany eed11ded30
server/.../safetensors: fix offsets and include all model parts (#9427) 1 year ago
Michael Yang b42aba40ed cuda: enable flash attention 1 year ago
王贺 25885e5335
docs: Add 1Panel to Community Integrations (#9312) 1 year ago
Jeffrey Morgan 98d44fa39d
llama: add phi4 mini support (#9403) 1 year ago
Blake Mizerany 2099e2d267
CONTRIBUTING: provide clarity on good commit messages, and bad (#9405) 1 year ago
Bruce MacDonald 0c1041ad85
runner: default to greedy sampler for performance (#9407) 1 year ago
Parth Sareen c245b0406f
sample: remove transforms from greedy sampling (#9377) 1 year ago
Michael Yang 8b194b7520 kvcache: update tests 1 year ago
Michael Yang 3e8b8a1933 ml: update Context.Forward interface 1 year ago
Blake Mizerany 41dc280491
server/internal/registry: implement CloseNotify and Flush (for now) (#9402) 1 year ago
Michael Yang 53d2990d9b model: add bos token if configured 1 year ago
Jesse Gross e185c08ad9 go.mod: Use full version for go 1.24.0 1 year ago
Blake Mizerany 2412adf42b
server/internal: replace model delete API with new registry handler. (#9347) 1 year ago
Steven Hartland be2ac1ed93
docs: fix api examples link (#9360) 1 year ago
Eries Trisnadi dc13813a03
server: allow vscode-file origins (#9313) 1 year ago
Michael Yang d6af13efed runner: simplify tensor split parsing 1 year ago
Michael Yang a59f665235 ml/backend/ggml: fix debug logging 1 year ago
Daniel Hiltgen 688925aca9
Windows ARM build (#9120) 1 year ago
Blake Mizerany 76e903cf9d
.github/workflows: swap order of go test and golangci-lint (#9389) 1 year ago
Jeffrey Morgan a5272130c4
ml/backend/ggml: follow on fixes after updating vendored code (#9388) 1 year ago
Jeffrey Morgan d7d7e99662
llama: update llama.cpp vendor code to commit d7cfe1ff (#9356) 1 year ago
Gordon Kamer 2db96c18e7
readme: add Nichey to community integrations (#9370) 1 year ago
Daniel Hiltgen e12af460ed
Add cuda Blackwell architecture for v12 (#9350) 1 year ago
Jeffrey Morgan 3ad4bc8afe
llama: removed unused 'vendoring' file (#9351) 1 year ago
Blake Mizerany 0d694793f2
.github: always run tests, and other helpful fixes (#9348) 1 year ago
Daniel Hiltgen e91ae3d47d
Update ROCm (6.3 linux, 6.2 windows) and CUDA v12.8 (#9304) 1 year ago
José Pekkarinen 6ecd7f64ba
docker: upgrade rocm to 6.3.3 (#8211) 1 year ago
Chuanhui Liu 888855675e
docs: rocm install link (#9346) 1 year ago
Michael Yang b16367b4b2 fix: add back bf16 support 1 year ago
Pavol Rusnak a499390648
build: support Compute Capability 5.0, 5.2 and 5.3 for CUDA 12.x (#8567) 1 year ago
frob 4df98f3eb5
Move cgroups fix out of AMD section. (#9072) 1 year ago
Blake Mizerany 348b3e0983
server/internal: copy bmizerany/ollama-go to internal package (#9294) 1 year ago
Parth Sareen 0b7e1676eb
sample: add sampling package for new engine (#8410) 1 year ago
Parth Sareen 314573bfe8
config: allow setting context length through env var (#8938) 1 year ago
Blake Mizerany 4604b10306
go.mod: bump to go1.24 (#9242) 1 year ago